The C-VCE framework integrates a concept bottleneck layer directly into diffusion models to generate visual counterfactuals. This architecture replaces fragile external classifiers with human-interpretable features to explain model predictions. It removes the need for noise-robust classifiers during image editing. Practitioners gain a more stable method for auditing vision models in safety-critical medical domains.