What was announced
Perplexity has released pplx-embed-v2-late, a family of two embedding models built on a ColBERT-style late-interaction approach. Instead of compressing a document into a single vector, the models keep a 128-dim vector for every token.
There are two sizes:
- 0.6B: designed to run on laptops, edge devices or small GPUs, and intended as a live query encoder.
- 9B: aimed at datacenters or high-memory GPUs, and intended for building the document index.
What it can search
Both models retrieve text, images and rendered PDF pages in one shared embedding space. Pages are encoded as images, so no OCR step is needed. The models are on Hugging Face under the MIT license, allowing commercial use. A hosted Perplexity API endpoint is planned but not live yet.
Headline results
According to Perplexity's own numbers, the 9B model reaches 92.4% accuracy on MADQA, an agentic PDF question-answering benchmark, while the 0.6B model scores 90.1%. The company says a 9B index can be searched with 0.6B queries, recovering roughly half the text quality gap at 0.6B query cost.
Limitations
- Storing one vector per token means index size grows with document length.
- It is not first on ViDoRe v3 image retrieval; Tencent's EVIE scores higher.
- A single input cannot mix text and images.
- All scores are self-reported and the technical report is not out yet.



