The C-VCE framework integrates a concept bottleneck layer directly into diffusion models to generate visual counterfactuals. This replaces fragile external classifiers with human-interpretable features. By removing the need for separate noise-robust classifiers, the system provides more stable explanations for safety-critical vision tasks. Practitioners gain a more reliable way to audit model predictions.