Researchers uncover that large language models can self‑assess meaning‑level confidence without explicit training, a finding that reshapes how AI systems judge their own answers. By applying a sampling‑based semantic calibration, the study demonstrates robust confidence estimates across open‑domain question answering. This insight could improve trust, safety, and deployment of AI in global applications.
The Signal
It originates from Apple Machine Learning Research.