The C-VCE framework integrates a classifier directly into a diffusion model via a concept bottleneck layer. This replaces fragile external classifiers that often fail on noisy images. By guiding edits through human-interpretable features, the method provides more robust counterfactual explanations. Practitioners in safety-critical fields like medicine can now better audit vision model predictions.