The GH-ESD framework identifies systematic failures in instance-level vision tasks by targeting spatially grounded patterns. Unlike previous methods that rely on simple attribute clusters, this approach isolates error slices in object detection and segmentation. Apple researchers use these hypotheses to pinpoint specific robustness gaps. This allows developers to fix precise visual failures.