Large language models have advanced rapidly, yet they still fail to master video games. Even when a few, like Gemini 2.5 Pro, beat simple titles, they do so slowly, erratically, and with custom tooling.
The Signal
Researchers at Modl.ai and other labs highlight that gaming remains a hard benchmark for current AI, underscoring gaps in real‑world reasoning and interaction.