The LFM2.5 model now offers Q4_0 checkpoints created via quantization-aware distillation. This method minimizes precision loss compared to standard post-training quantization. It allows smaller hardware footprints without sacrificing significant performance. Developers can now deploy these optimized weights to reduce VRAM usage during local inference tasks.