The new DiG-bench benchmark uses 70 games to measure how AI systems infer unwritten environmental rules through exploration. It tests whether models like Fable possess creative intuition or simply follow fed data. This shift toward discovery-based evaluation helps researchers quantify curiosity. Practitioners can use these metrics to build more autonomous, adaptive agents.