A new benchmark called DiG-bench uses 70 games to measure how AI systems infer unwritten environmental rules through exploration. This test shifts focus from static data to active curiosity and creative intuition. Researchers use these interactive systems to map discovery capabilities. Practitioners can now quantify how well models learn without explicit instructions.