Schmidt Sciences calls for proposals to build interpretability tools that spot and curb deceptive behavior in large language models. The pilot seeks methods that detect when LLMs give misleading or harmful advice and steer reasoning away from such outputs. Success could unlock a larger investment in trustworthy AI practices.
The Signal
Funding will target real‑world applications and measurable risk mitigation.