From the source
Perplexity introduced pplx-embed-v2-late, a family of late-interaction embedding models that combine multi-vector representations, support for both text and image modalities, and a shared embedding space across model sizes (0.6B and 9B).
The models achieve state-of-the-art results on vision and text retrieval benchmarks, with the 0.6B variant matching models five times larger on the ViDoRe benchmark, and are publicly available on Hugging Face.