Perplexity AI releases pplx-embed-v2-late models for multimodal search

1 hour ago 1



Perplexity AI has released two new embedding models that read pages more like people do: words, pictures, tables and all. The models, called pplx-embed-v2-late, come in 0.6B and 9B parameter versions and are now publicly available on Hugging Face. Announced on October 7, 2026, the pair handles retrieval across text, images and visual documents such as PDFs and slide decks. What Perplexity actually shipped Perplexity’s new models use a ColBERT-style late-interaction architecture that keeps a separate 128-dimensional vector for each token. Rather than one summary per document, the system keeps a running set of notes on every piece. Matching is handled with MaxSim scoring. Each part of a query looks for its best match inside a document, and those best matches drive the final score. That shared space enables the release’s most practical feature. A corpus indexed with the larger 9B model can be searched using the lightweight 0.6B model, with testing showing this cross-model querying worked efficiently at no additional cost. Translated for the people paying cloud bills: you can index once with the heavyweight model, then answer live queries with the smaller, faster one. The numbers behin...

Read Entire Article