The DiG-bench benchmark uses 70 games to measure how AI systems infer unwritten environmental rules through exploration. This test moves beyond static datasets to evaluate creative intuition and curiosity. Researchers use these interactive systems to map discovery capabilities. Practitioners can now quantify whether a model truly understands a system or simply follows fed instructions.