Skip to content

pplx-embed: Perplexity's Embedding Model Family and How to Choose Between Its Three Lines

Oct 9, 20261 min
TL;DRpplx-embed is the embedding family Perplexity has been releasing since February 2026, in three lines: v1 single-vector (0.6B/4B, API at $0.004/$0.03 per 1M tokens), context embeddings (v1 plus a 9B v2 preview), and v2-late multi-vector (0.6B/9B, searches PDF page images directly). They solve different problems; v2 does not replace v1.

🌏 中文版

On October 7, 2026, Perplexity released pplx-embed-v2-late, a week after pplx-embed-v2-context-9b-preview, and the original pplx-embed-v1 dates from February. That is three batches and seven checkpoints in eight months, all named pplx-embed, but "v2" is not an upgrade of "v1". This is the twenty-fifth post in the "AI model family" series. It splits the family into three lines and answers one question: which one fits your documents.

For how to read the benchmark numbers quoted here, see the AI model evaluation sources guide. For where models sit overall, see the AI model landscape overview. The v2-late details have their own post: pplx-embed-v2-late explained.

Timeline

DateModelKey facts
2026-02-11Technical reportarXiv:2602.11151 appears before the models
2026-02-26pplx-embed-v1, pplx-embed-context-v10.6B and 4B of each, four checkpoints; MIT license, API on day one
2026-09-30pplx-embed-v2-context-9b-preview9B contextual preview, built with turbopuffer; weights on Hugging Face, no API yet
2026-10-07pplx-embed-v2-late-0.6b / 9bMulti-vector, text plus images; MIT weights, API planned

v2 currently has only the context and late lines. There is no v2 dense model yet: the late announcement says Perplexity is "currently training the next generation of our dense embedding models" and will roll them out on the API progressively. For ordinary semantic search today, v1 is what you can actually use.

Three lines, three problems

v1 (dense)context (v1 / v2 preview)v2-late
Output per inputOne vectorOne vector per chunk, aware of the whole documentOne 128-dim vector per token
Sizes0.6B (1024-dim), 4B (2560-dim)v1: 0.6B, 4B; v2: 9B (2048-dim, truncatable to 1024)0.6B, 9B
InputTextTextText, page images
QuantizationNative INT8 / binaryv1: INT8 / binary; v2 preview: INT8Not stated
Max length32Kv2 evaluated at up to 32,768 tokens per passNot published
ScoringCosineCosineMaxSim
API (checked 2026-10-09)Yesv1 yes, v2 preview noNo, planned

Prices from the Perplexity API docs, checked 2026-10-09, per 1M tokens:

ModelPrice
pplx-embed-v1-0.6b$0.004
pplx-embed-v1-4b$0.03
pplx-embed-context-v1-0.6b$0.008
pplx-embed-context-v1-4b$0.05

v1: turning a decoder into an encoder with diffusion pretraining

v1 starts from an unusual choice. Most strong embedding models sit on decoder-only language models, where a token only sees what precedes it. Perplexity starts from Qwen3 (0.6B and 4B), removes the causal mask, and continues pretraining with a diffusion denoising objective on about 250 billion tokens across 30 languages, producing a bidirectional encoder before contrastive training. Its ablations credit this with roughly one percentage point on retrieval.

Three contrastive stages follow: pair training, contextual training (which yields the context models) and triplet training with mined hard negatives. The final v1 model merges the contextual and triplet checkpoints by spherical linear interpolation. Quantization is part of training rather than a post-hoc step: embeddings run as INT8 throughout with straight-through estimation, cutting storage 4x against FP32; binary cuts it 32x, with a drop under 1.6 points at 4B and 2 to 4 points at 0.6B.

Another trade-off: no instruction prefix. The model card says prefixes can add 2% to 3% on benchmarks but make indexing brittle when index-time and query-time prefixes drift apart, so they were left out.

Reported results (all self-published):

ItemNumberComparison
MTEB(Multilingual, v2) retrieval, 4B INT869.66%Qwen3-Embedding-4B 69.60%, gemini-embedding-001 67.71%
ConTEB, context-v1-4B INT881.96%voyage-context-3 79.45%, Anthropic Contextual 72.4%
PPLXQuery2Doc (30M docs), 4BRecall@1000 91.7%Qwen3-Embedding-4B 88.6%

On MTEB the gap is 0.06 points, effectively a tie. v1's pitch is parity at INT8, which makes storage cheap.

The context line: from the "gold passage" to answer plus evidence

This line targets the old chunking problem: once a passage is cut out, its references, headings and definitions are gone. The method is late chunking: encode the whole document in one pass, then pool within chunk boundaries, so every chunk vector has seen the full text.

The v2 preview changes the supervision. The usual recipe has an LLM label one "gold passage" and treats every other chunk as a negative, including chunks that carry supporting evidence. The preview instead uses Perplexity's query-aware context compression model as a teacher: it scores each document token for relevance, those scores are aggregated into chunk scores, and the embedding model learns to retrieve the answer chunk together with its support. It starts from an in-house 9B ColBERT model, trains on about 430 datasets in more than 50 languages, and ships as a model soup averaged over several checkpoints.

This is the line whose evaluation needs the most careful reading:

  • The headline results are on context-bench, held privately by turbopuffer, which only evaluates submitted models. That limits training contamination but nobody outside can reproduce it.
  • On context-bench at K=10 the 9B preview reaches 45.5% answer recall and 40.6% evidence recall, 14.4 and 5.0 points above voyage-context-4. The absolute levels are modest: document-level Document@10 is 61.6%, so the benchmark is hard for every model.
  • It has the best average on the public ConTEB but not every task: it loses NarrativeQA to the v1 4B model and COVID-QA to Nemotron. Perplexity notes COVID-QA rewards surface lexical matching, so context helps less there.
  • On query-to-document retrieval across domains, its average is slightly below voyage-context-4.
  • The v2 post also says the v1 context-4B was trained only on the ConTEB training set and transfers poorly to other domains. That comes from Perplexity itself, and it means v1 context's 81.96% on ConTEB should not be extrapolated to your domain.

One practical storage figure: the preview emits 1024-dim INT8 vectors at 1 KB each and slightly beats voyage-context-4's 2048-dim float32 (8 KB) on average chunk-retrieval score. That counts vectors only, not the rest of the index.

Usage notes from the model card:

  • It is a preview; weights and interface may change without backward compatibility, so do not mix preview embeddings with a future release.
  • Use encode_queries for queries and encode for document chunks; they use different prefixes and the wrong call "silently degrades retrieval quality".
  • Embeddings are unnormalized INT8; compare with cosine similarity.
  • It needs transformers>=5.4.0 and trust_remote_code=True.
  • The model card does not state a license. Secondary coverage says MIT, which I did not confirm on a primary page; check the license field on Hugging Face before use.

v2-late: a vector per token, and images

The architecture and evaluation of v2-late are covered in the dedicated post; here is only its place in the family. It is the one multi-vector line: a Qwen3.5 base, an 18B teacher distilled into 9B and 0.6B students, and a shared embedding space so a 9B index can be queried by the 0.6B. Training data is 186 million query–document pairs from 594 datasets in 46 languages.

The trade-off to remember: it has the highest storage cost of the three lines, in exchange for page-image retrieval and agent-style accuracy. The MADQA 92.4% is a self-run test with a Gemini 3.5 Flash agent; the dedicated post has the conditions.

Selection matrix

Your situationPickWhy
General semantic search, text only, need it nowv1 4B; 0.6B on a tight budgetOnly dense line with an API; INT8 storage is small
Chunked long documents with many references and definitionsv1 context (has API), or evaluate the v2 previewLate chunking keeps document context; the preview is not for production
Need both the answer and supporting passages for an agent or a reviewerEvaluate the v2-context previewThat is its training goal, but the benchmark is private
PDFs with tables and charts where OCR is the bottleneckEvaluate v2-lateOnly line that searches page images directly
Tight storage, vector DB without multi-vector supportv1v2-late index grows with token count
Must use an API, no self-hostingv1 or v1 contextBoth v2 lines are weights-only for now

Related reading on this site: BGE-M3 embedding model selection and ColPali visual document retrieval.

Bottom line

Under one name, pplx-embed is three bets. v1 bets that native INT8 and no instruction prefix make single vectors cheap enough. The context line bets that chunks must carry document context and learn evidence, not just answers. The late line bets that a pricier index buys image retrieval and agent accuracy. What they share is Perplexity's own search traffic as the yardstick, which is close to real queries but means the most prominent benchmarks (PPLXQuery, PPLX-Q2I, context-bench) are not available to outsiders.

Three things to watch: when a v2 dense model ships and whether it reaches the API; when the context preview becomes a final release and whether anyone outside turbopuffer replicates the results; and whether the late technical report lets others reproduce 92.4%. Until then, production workloads belong on v1 with its API and stable versions, and the other two lines are evaluation candidates.

References