The AFM 3 Core Advanced model powers Siri's new expressive voices using a decoupled temporal depth diffusion transformer. This architecture converts semantic audio tokens into high-fidelity speech within the strict memory limits of the Apple Matrix Coprocessor. It utilizes residual vector quantization to ensure real-time performance. This optimizes high-quality audio generation for constrained edge hardware.