• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

Perplexity introduces contextual embedding preview with compatibility warning

Perplexity claims leading Answer and Evidence retrieval on context-bench, while its earlier model leads two Document measures. The preview’s embeddings may be incompatible with future releases.

Aravind SrinivasAS
BoBO
turbopufferTU
12 Sources, 10d ago, first seen 10d ago

TLDR

Perplexity’s contextual embedding preview encodes document chunks together and trains on relevance signals meant to retrieve answers and supporting context. The company claims leading Answer and Evidence retrieval on context-bench, but its earlier model is slightly ahead on Document@3 and Document@5. The model card warns that future versions may break compatibility and advises against mixing preview embeddings with later releases.

Combined views

152.6K

12 Sources, first seen 10d ago

1.4K likes53 comments686 saves99 reposts

Combined views

152.6K

12 Sources, first seen 10d ago

1.4K likes53 comments686 saves99 reposts

Perplexity has introduced pplx-embed-v2-context-9b-preview, a contextual embedding model that processes document chunks together so each chunk’s representation reflects surrounding context. The company announced the preview on Sept. 30.

Featured Source

Its announcement thread describes the aim as retaining context that can be lost when retrieval systems split a long document into isolated chunks.

Training beyond a single answer chunk

Perplexity describes a training signal drawn from its query-aware context compression model. That model scores document tokens against a query; those scores are aggregated into chunk-level targets. The company says this trains the embedder to retrieve both answer chunks and supporting material.

On turbopuffer’s privately held context-bench, Perplexity claims the preview leads Answer and Evidence retrieval at every cutoff. The same update qualifies the comparison: its earlier context v1 4B model is slightly ahead on Document@3 and Document@5. The claimed lead therefore does not cover every document-retrieval measure.

A preview with compatibility limits

The model card documents 2,048-dimensional INT8 embeddings and training at both 1,024 and 2,048 dimensions. It also requires different encoding methods for queries and document chunks, warning that using the document method for queries degrades retrieval quality.

Perplexity warns that this is a preview rather than a final model. Weights, embeddings and the interface may change without backward compatibility. Its model card advises against mixing embeddings produced by this preview with those from a future release.

Sentiment

Positive84.9%15.1%Negative

Based on 73 sentiment-bearing replies from 64 accounts across 6 conversations.

Sentiment

Positive84.9%15.1%Negative

Based on 73 sentiment-bearing replies from 64 accounts across 6 conversations.

Related

Perplexity releases updated open-weights decision model, with benchmark and price claims attached

Perplexity says its new pplx-decider-v1.1-27b model costs $0.02 per million input tokens, half the price of v1. OpenRouter later said the model was also available there with text, JSON, or image inputs and free output.

Denis Bykov asked disappointed Perplexity API users to DM him

In an X post, Denis Bykov asked people who had tried Perplexity’s APIs to message him if they were disappointed by anything. In a reply, CEO Aravind Srinivas said, “We would love to hear how we can improve our developer platform!”

Perplexity introduces a Decisions API built around pplx-decider-v1-27b

Perplexity introduced a Decisions API built around pplx-decider-v1-27b, a multimodal model that returns probabilities over fixed answers instead of text, with weights available on Hugging Face.

13 Sources

huggingfaceperplexity-ai/pplx-embed-v2-context-9b-preview · Hugging Face
Perplexity@perplexity_aiWe built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage10d
Bo@bo_wangboin the past months we worked closely with tpuf team, iterating our models beyond web search. During the journey, we built a new way of training contextual embedding models, which connects our previous work, the context pruner. We're sharing how we train the model, and how it performs on @turbopuffer 's private bench. great work from @ESL_Sarah , Markus, @antoine_chaffin , Louis, Max and special thanks to @n0riskn0r3ward , who helps evaluate, iterate again and again 🐐. We'll work closely together to continue improve it, and eventually offer it to all.10d
Aravind Srinivas@AravSrinivasWe’re open sourcing our state-of-the-art contextual embedding models, which perform best in turbopuffer’s context-bench.10d
Denis Yarats@denisyaratswe developed a new way to train contextual embedding models. instead of embedding each chunk on its own, the model encodes the whole document once and pools chunk vectors afterward, so every chunk sees the full document. the new part is the training signal: instead of one labeled "gold" chunk per query, we distill relevance from our context compression model, which scores every document token against the query. this teaches the model to retrieve both the answer and the context that supports it. pplx-embed-v2-context-9b-preview sets a new SOTA on ConTEB and on turbopuffer's new private context-bench, which it was evaluated on blind: +14.4 points answer recall@10 over voyage-context-4. and at 1024 dims in int8 (1KB per vector), it still beats voyage-context-4 at 8KB per vector.10d
turbopuffer@turbopufferwe maintain several internal benchmarks to guide our customers toward better search relevance @perplexity's new model tops context-bench, our internal contextual embedding benchmark, and boosts document recall@10 by 50%+ over traditional SOTA embedding models10d
Mikhail Parakhin@MParakhinPrediction: "in context of ..." encoding will be more and more popular. Back in the day I was making fun of the rudimentary "split into paragraphs" approach for RAG, popularized in one famous tutorial - it is getting better now :-) Turbopuffer founder is Shopify alumni, btw10d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Perplexity

    13 Sources

    huggingfaceperplexity-ai/pplx-embed-v2-context-9b-preview · Hugging Face
    Perplexity@perplexity_aiWe built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage10d
    Bo@bo_wangboin the past months we worked closely with tpuf team, iterating our models beyond web search. During the journey, we built a new way of training contextual embedding models, which connects our previous work, the context pruner. We're sharing how we train the model, and how it performs on @turbopuffer 's private bench. great work from @ESL_Sarah , Markus, @antoine_chaffin , Louis, Max and special thanks to @n0riskn0r3ward , who helps evaluate, iterate again and again 🐐. We'll work closely together to continue improve it, and eventually offer it to all.10d
    Aravind Srinivas@AravSrinivasWe’re open sourcing our state-of-the-art contextual embedding models, which perform best in turbopuffer’s context-bench.10d
    Denis Yarats@denisyaratswe developed a new way to train contextual embedding models. instead of embedding each chunk on its own, the model encodes the whole document once and pools chunk vectors afterward, so every chunk sees the full document. the new part is the training signal: instead of one labeled "gold" chunk per query, we distill relevance from our context compression model, which scores every document token against the query. this teaches the model to retrieve both the answer and the context that supports it. pplx-embed-v2-context-9b-preview sets a new SOTA on ConTEB and on turbopuffer's new private context-bench, which it was evaluated on blind: +14.4 points answer recall@10 over voyage-context-4. and at 1024 dims in int8 (1KB per vector), it still beats voyage-context-4 at 8KB per vector.10d
    turbopuffer@turbopufferwe maintain several internal benchmarks to guide our customers toward better search relevance @perplexity's new model tops context-bench, our internal contextual embedding benchmark, and boosts document recall@10 by 50%+ over traditional SOTA embedding models10d
    Mikhail Parakhin@MParakhinPrediction: "in context of ..." encoding will be more and more popular. Back in the day I was making fun of the rudimentary "split into paragraphs" approach for RAG, popularized in one famous tutorial - it is getting better now :-) Turbopuffer founder is Shopify alumni, btw10d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet