The LFM2.5 series now includes Q4_0 checkpoints created via quantization-aware distillation. This technique reduces model size while minimizing the typical accuracy loss found in standard post-training quantization. Developers can now deploy these smaller weights on consumer hardware. It is an incremental optimization for efficiency rather than a leap in capability.