GH-ESD identifies systematic failures in instance-level vision tasks by targeting spatially grounded visual patterns. Unlike previous methods that rely on simple clusters, this approach discovers error slices through hypothesis-driven discovery. It allows researchers to pinpoint exactly why object detection models fail in specific contexts. This improves how developers debug vision robustness.