The AFM 3 Core Advanced model uses a decoupled temporal depth diffusion transformer to synthesize high-fidelity speech on-device. This architecture converts semantic audio tokens into residual vector quantization representations. It optimizes for the Apple Matrix Coprocessor memory budget. The design enables real-time, configurable voice synthesis without relying on cloud compute.