A new benchmark of 70 games called DiG-bench measures how AI systems infer unwritten environmental rules through exploration. The tests evaluate creative intuition and curiosity rather than reliance on provided data. Fable showed early promise in these discovery tasks. This provides a controlled framework for researchers to quantify autonomous learning capabilities.