The DiG-bench benchmark uses 70 interactive games to measure how AI systems infer unwritten environmental rules through exploration. It tests curiosity-driven discovery rather than reliance on provided instructions. This framework helps researchers quantify the creative intuition of models like Fable. Practitioners can use these metrics to evaluate agentic reasoning in unstructured environments.