A new benchmark pits large language models against each other in a 1v1 real‑time strategy game, forcing them to write code that controls units. By treating code generation as a competitive, dynamic task, researchers from OpenAI and DeepMind can better gauge model reasoning, planning, and adaptability.
The Signal
The format promises rapid, reproducible insights that could accelerate AI safety and deployment worldwide.