pplx-embed-v2-late: An Open Late-Interaction Model That Searches PDF Pages Without OCR
On October 7, Perplexity open-sourced two ColBERT-style multimodal embedding models (0.6B and 9B, MIT license) that keep one 128-dim vector per token, score with MaxSim, and search PDF page images directly. The 9B scores 92.4% on MADQA (self-reported, with a Gemini 3.5 Flash agent); the cost is an index that grows with document length.