Apple researchers introduced GH-ESD to identify systematic failures in instance-level vision tasks. Unlike previous methods using simple clusters, this approach targets spatially grounded visual patterns and contextual relations. It specifically solves the difficulty of finding error slices in object detection and segmentation. Practitioners can now pinpoint precise semantic weaknesses to improve model robustness.