AIREITER

pplx-embed-v2-late API Pricing and PDF Retrieval Setup

Last Updated: 2026-10-07 19:14:32

If you are budgeting a PDF search system around pplx-embed-v2-late, the important answer is easy to miss: Perplexity has released the weights, but it has not published a v2-late API price or listed the model in its public Embeddings API catalog. The usable path today is self-hosted multimodal retrieval, with the current v1 API prices serving only as a comparison point.

The pricing answer: v2-late is not on the public Embeddings API rate card

As of October 7, 2026, the official Embeddings API quickstart lists four v1 models. It does not list pplx-embed-v2-late-0.6b or pplx-embed-v2-late-9b, so there is no defensible per-token API estimate for v2-late yet.

Perplexity model currently listed in the API docsPrice per 1M tokensIntended input
pplx-embed-v1-0.6b$0.004Independent text, queries, sentences
pplx-embed-v1-4b$0.030Independent text, queries, sentences
pplx-embed-context-v1-0.6b$0.008Related document chunks
pplx-embed-context-v1-4b$0.050Related document chunks

These are pay-as-you-go API rates, not prices for the late-interaction family. The Perplexity release announcement says late-interaction, dense, and contextual embeddings will be rolled out progressively on the API Platform. That is a rollout statement, not an active v2-late endpoint or a price commitment.

For a purchase decision, split the budget into two lines:

  1. Managed API spend: available for the v1 models above; no v2-late rate is published.
  2. Self-hosting spend: GPU time, page rendering, model storage, token-vector index storage, and query serving for v2-late.

Do not multiply a v1 price by a PDF page count and call the result a v2-late quote. The models use different representations and the v1 API is text embedding, not the documented rendered-page workflow.

What pplx-embed-v2-late actually ships

Perplexity publishes two late-interaction checkpoints: pplx-embed-v2-late-0.6b and pplx-embed-v2-late-9b. The 9B model card reports 340M active parameters for the smaller model and 7.4B for the larger one. Both output 128-dimensional vectors per token and use MaxSim rather than reducing a page to one vector.

ModelActive parametersViDoRe v3 image nDCG@10ViDoRe v3 Markdown nDCG@10Practical role
pplx-embed-v2-late-0.6b340M62.3%61.2%Lower-weight query or smaller deployment
pplx-embed-v2-late-9b7.4B65.2%64.7%Higher-quality index and retrieval model

The benchmark figures are the model card's results, not an independent PDF test: 9B leads by 2.9 percentage points on image retrieval and 3.5 points on Markdown retrieval, with about 21.8 times as many active parameters. Both checkpoints are MIT-licensed on Hugging Face.

The shared embedding space is the key deployment detail: Perplexity says a 9B-built index can be queried with the 0.6B model. Use 9B for offline document encoding and 0.6B for queries only after validating cross-model recall; this does not remove 9B index storage.

A workable PDF retrieval setup

The v2-late workflow treats each rendered PDF page as an image document. A text query can then match a page's words, table structure, chart, or layout without requiring OCR as the primary retrieval representation. This is the same visual-document pattern described in Sentence Transformers' visual retrieval documentation.

"OCR-free" means OCR is not the retrieval signal; extracted text is still useful for filtering, citations, accessibility, and fallback search.

1. Render pages and preserve metadata

Render every page to an RGB image at a stable resolution, then store a record beside it:

FieldExample
document_idcontract-2026-04
page_number17
image_pathpages/contract-2026-04/017.png
source_uriInternal PDF object URL
text_fallbackOptional extracted text

Keep document_id and page_number in the retrieval record, not inside the image filename alone. After a relevant page is found, fetch adjacent pages from the same document because a table, footnote, or definition often crosses a page boundary.

2. Install the compatible encoder

The 9B model card requires recent libraries:

pip install "sentence-transformers>=6.0.0" "transformers>=5.4.0" pillow

The authored example uses MultiVectorEncoder, and the model card selects CUDA for loading the 9B checkpoint:

from PIL import Image
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder(
    "perplexity-ai/pplx-embed-v2-late-9b",
    device="cuda",
)

Use the 0.6B identifier when the larger checkpoint does not fit the available serving hardware. The model card gives no official VRAM minimum, throughput table, or latency guarantee, so measure your page resolution, batch size, and GPU before committing to capacity.

3. Encode text queries and page images separately

The model requires asymmetric calls. Query text goes through encode_query; rendered pages go through encode_document:

query_embeddings = model.encode_query([
    "Which clause governs termination after a material breach?"
])

page = Image.open("pages/contract-2026-04/017.png").convert("RGB")
page_embeddings = model.encode_document([page])

scores = model.similarity(query_embeddings, page_embeddings)
print(scores)

Do not put text and image documents in one mixed encoding batch. The model card specifically calls out separate homogeneous inputs and the [Q] / [D] marker configuration expected by this checkpoint. model.similarity() applies MaxSim over the token-level representations.

For a real collection, encode pages offline, persist the multi-vector representation in a late-interaction index, and keep the page metadata in a sidecar store. A small collection can use exhaustive scoring. At larger scale, use a system that supports MaxSim or use a dense first-stage retriever followed by v2-late reranking of a controlled candidate set.

4. Retrieve pages, then expand the evidence window

A page-level hit should normally return:

  1. The matched page and its score.
  2. The document ID and source link.
  3. One or two neighboring pages from the same document.
  4. The page image and any optional extracted text for citation.

This avoids turning a visually accurate page match into an incomplete answer when the definition starts on page 16 and the table continues on page 17. It also makes the output inspectable: a user can see the chart or table that produced the match instead of trusting a hidden OCR transformation.

The cost model beyond API tokens

There is no published v2-late API price to compare with the four v1 rates. The operational cost is therefore dominated by deployment choices that the model card leaves unpriced.

Cost driverWhat is confirmedPlanning implication
Model weightsThe 9B repository shows about 33.6 GB and F32 tensors on Hugging FaceWeight storage and loading are material before indexing begins
RepresentationOne 128-dimensional vector per token, scored with MaxSimA page produces many vectors, not one dense vector
Indexing9B can build an index queryable by 0.6BPut higher compute in an offline job if query volume is high
RetrievalLate interaction compares query tokens with document tokensUse a supported MaxSim index or limit candidates before rescoring
API billingNo v2-late rate is publishedDo not forecast managed API spend yet

The Hugging Face late-interaction guide gives a useful scale reference from another model: a 4,874-passage example produced 608,414 token vectors and 311.5 MB of raw float32 storage, while a compressed PLAID index used 92 MB. Those numbers are not a v2-late estimate, but they show why “128 dimensions” does not equal “small index.” Token count is the important multiplier.

Indexing throughput also deserves a benchmark on your own hardware. In a real-user LocalLLaMA report, pplx-embed-v1-4b took about 45 minutes per 10,000 vectors versus 6 minutes for Qwen3-Embedding-4B on an A100 80GB. That report concerns v1, not v2-late, so it is a warning about measuring Perplexity embedding throughput, not a v2-late performance claim.

“I think it might be because pplx embed uses bidirectional attention rather than standard masked attention.” — u/Velocita84, r/LocalLLaMA

Which deployment path should you choose?

RequirementBest current pathWhy
Cheap text-only RAG with a managed endpointPerplexity v1 APIPublished prices run from $0.004 to $0.05 per 1M tokens
Charts, tables, scanned pages, and layout matterSelf-host pplx-embed-v2-lateThe documented workflow searches rendered pages directly
Large corpus with frequent queries9B offline index plus 0.6B query encoder, or dense-first plus v2-late rerankingSeparates indexing quality from query-time compute
Small prototype or hardware-constrained test0.6B checkpoint on a representative page sampleLower active-parameter count, but still benchmark page encoding and storage
Managed v2-late endpoint is a hard requirementWait for an official API model ID and rate cardNeither is present in the current public embedding documentation

My recommendation is to prototype the 0.6B and 9B cross-model path on 100 to 500 representative pages before building a full index. Include scanned pages, tables, multi-column layouts, and pages whose answer spans a boundary. Record recall at your target k, page-encoding throughput, raw and compressed index size, and query latency. That evidence is more useful than transferring the v1 token price to a model that is not yet sold through that API.

pplx-embed-v2-late PDF retrieval FAQ

Does pplx-embed-v2-late have an API price?

Not in the public Perplexity Embeddings API documentation checked for this guide. The published $0.004 to $0.05 per million-token prices apply to v1 standard and contextualized models.

Is pplx-embed-v2-late officially released?

Yes. Perplexity publishes 0.6B and 9B open-weight checkpoints on Hugging Face. Weight release and managed API availability are separate milestones.

Does PDF retrieval require OCR?

No, not for the visual retrieval signal. Render each page as an image and encode it as a document. OCR or extracted text remains useful for filtering, citations, accessibility, and fallback search.

Can the 0.6B model query a 9B-built index?

Perplexity’s model card says the two models share an embedding space and support that arrangement. Measure quality on your corpus because the card does not publish a cross-model retrieval delta.

Can text and image pages share one batch?

No. The model card says mixed text-and-image inputs are not supported in the same encoding batch. Keep text and image encoding calls homogeneous.

Do I still need a reranker?

Not automatically. MaxSim is already the late-interaction scoring method, but a dense first-stage retriever plus v2-late reranking can be more practical than scanning every page-token vector in a large corpus.

What is the exact storage cost per PDF page?

Perplexity does not publish a v2-late page-level storage calculator. Estimate it from the number of retained page tokens, vector precision, metadata, and index compression, then validate with a representative sample.

Should I choose 0.6B or 9B?

Use 9B when offline indexing quality is the priority and you can afford the model and indexing job. Use 0.6B for a smaller deployment or query encoder, including the documented shared-space setup against a 9B index. The benchmark gap is measurable, but the model card does not provide a universal quality or latency rule.

The go/no-go check is short: if you need a managed Perplexity endpoint with a known price today, v2-late is not ready for that requirement. If you can self-host and your PDFs contain information that OCR or text chunking loses, render a representative page set and benchmark the late-interaction pipeline before scaling it.