Apple’s new AMES framework unifies text, image, and video queries in a single shared space, enabling late‑interaction retrieval without redesigning existing engines. By embedding tokens, patches, and frames with multi‑vector encoders, AMES supports cross‑modal search in production.
The Signal
The two‑stage pipeline first performs parallel token‑level ANN search, then refines results across modalities.