The DiG-bench benchmark uses 70 games to measure how AI systems infer unwritten environmental rules through exploration. It moves beyond static datasets to test curiosity and creative intuition. Researchers can now quantify whether a model truly discovers logic or simply relies on pre-fed patterns. This provides a concrete metric for agentic discovery.