The dots.mocr model achieved 97.6 percent character accuracy across 2,000 historical book pages. Hugging Face and EleutherAI tested 14 open-source OCR tools to find cost-effective ways to digitize training sets. This accuracy suffices for language model training but fails scholarly standards. It lowers the cost of high-quality data ingestion to under $2 per thousand pages.