If you are budgeting a PDF search system around pplx-embed-v2-late, the important answer is easy to miss: Perplexity has released the weights, but it has not published a v2-late API price or listed the model in its public Embeddings API catalog. The usable path today is self-hosted multimodal retrieval, with the current v1 API prices serving only as a comparison point.
The pricing answer: v2-late is not on the public Embeddings API rate card
As of October 7, 2026, the official Embeddings API quickstart lists four v1 models. It does not list pplx-embed-v2-late-0.6b or pplx-embed-v2-late-9b, so there is no defensible per-token API estimate for v2-late yet.
| Perplexity model currently listed in the API docs | Price per 1M tokens | Intended input |
|---|---|---|
pplx-embed-v1-0.6b | $0.004 | Independent text, queries, sentences |
pplx-embed-v1-4b | $0.030 | Independent text, queries, sentences |
pplx-embed-context-v1-0.6b | $0.008 | Related document chunks |
pplx-embed-context-v1-4b | $0.050 | Related document chunks |
These are pay-as-you-go API rates, not prices for the late-interaction family. The Perplexity release announcement says late-interaction, dense, and contextual embeddings will be rolled out progressively on the API Platform. That is a rollout statement, not an active v2-late endpoint or a price commitment.
For a purchase decision, split the budget into two lines:
- Managed API spend: available for the v1 models above; no v2-late rate is published.
- Self-hosting spend: GPU time, page rendering, model storage, token-vector index storage, and query serving for v2-late.
Do not multiply a v1 price by a PDF page count and call the result a v2-late quote. The models use different representations and the v1 API is text embedding, not the documented rendered-page workflow.
What pplx-embed-v2-late actually ships
Perplexity publishes two late-interaction checkpoints: pplx-embed-v2-late-0.6b and pplx-embed-v2-late-9b. The 9B model card reports 340M active parameters for the smaller model and 7.4B for the larger one. Both output 128-dimensional vectors per token and use MaxSim rather than reducing a page to one vector.
| Model | Active parameters | ViDoRe v3 image nDCG@10 | ViDoRe v3 Markdown nDCG@10 | Practical role |
|---|---|---|---|---|
pplx-embed-v2-late-0.6b | 340M | 62.3% | 61.2% | Lower-weight query or smaller deployment |
pplx-embed-v2-late-9b | 7.4B | 65.2% | 64.7% | Higher-quality index and retrieval model |
The benchmark figures are the model card's results, not an independent PDF test: 9B leads by 2.9 percentage points on image retrieval and 3.5 points on Markdown retrieval, with about 21.8 times as many active parameters. Both checkpoints are MIT-licensed on Hugging Face.
The shared embedding space is the key deployment detail: Perplexity says a 9B-built index can be queried with the 0.6B model. Use 9B for offline document encoding and 0.6B for queries only after validating cross-model recall; this does not remove 9B index storage.
A workable PDF retrieval setup
The v2-late workflow treats each rendered PDF page as an image document. A text query can then match a page's words, table structure, chart, or layout without requiring OCR as the primary retrieval representation. This is the same visual-document pattern described in Sentence Transformers' visual retrieval documentation.
"OCR-free" means OCR is not the retrieval signal; extracted text is still useful for filtering, citations, accessibility, and fallback search.
1. Render pages and preserve metadata
Render every page to an RGB image at a stable resolution, then store a record beside it:
| Field | Example |
|---|---|
document_id | contract-2026-04 |
page_number | 17 |
image_path | pages/contract-2026-04/017.png |
source_uri | Internal PDF object URL |
text_fallback | Optional extracted text |
Keep document_id and page_number in the retrieval record, not inside the image filename alone. After a relevant page is found, fetch adjacent pages from the same document because a table, footnote, or definition often crosses a page boundary.
2. Install the compatible encoder
The 9B model card requires recent libraries:
pip install "sentence-transformers>=6.0.0" "transformers>=5.4.0" pillow
The authored example uses MultiVectorEncoder, and the model card selects CUDA for loading the 9B checkpoint:
from PIL import Image
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder(
"perplexity-ai/pplx-embed-v2-late-9b",
device="cuda",
)
Use the 0.6B identifier when the larger checkpoint does not fit the available serving hardware. The model card gives no official VRAM minimum, throughput table, or latency guarantee, so measure your page resolution, batch size, and GPU before committing to capacity.
3. Encode text queries and page images separately
The model requires asymmetric calls. Query text goes through encode_query; rendered pages go through encode_document:
query_embeddings = model.encode_query([
"Which clause governs termination after a material breach?"
])
page = Image.open("pages/contract-2026-04/017.png").convert("RGB")
page_embeddings = model.encode_document([page])
scores = model.similarity(query_embeddings, page_embeddings)
print(scores)
Do not put text and image documents in one mixed encoding batch. The model card specifically calls out separate homogeneous inputs and the [Q] / [D] marker configuration expected by this checkpoint. model.similarity() applies MaxSim over the token-level representations.
For a real collection, encode pages offline, persist the multi-vector representation in a late-interaction index, and keep the page metadata in a sidecar store. A small collection can use exhaustive scoring. At larger scale, use a system that supports MaxSim or use a dense first-stage retriever followed by v2-late reranking of a controlled candidate set.
4. Retrieve pages, then expand the evidence window
A page-level hit should normally return:
- The matched page and its score.
- The document ID and source link.
- One or two neighboring pages from the same document.
- The page image and any optional extracted text for citation.
This avoids turning a visually accurate page match into an incomplete answer when the definition starts on page 16 and the table continues on page 17. It also makes the output inspectable: a user can see the chart or table that produced the match instead of trusting a hidden OCR transformation.
The cost model beyond API tokens
There is no published v2-late API price to compare with the four v1 rates. The operational cost is therefore dominated by deployment choices that the model card leaves unpriced.
| Cost driver | What is confirmed | Planning implication |
|---|---|---|
| Model weights | The 9B repository shows about 33.6 GB and F32 tensors on Hugging Face | Weight storage and loading are material before indexing begins |
| Representation | One 128-dimensional vector per token, scored with MaxSim | A page produces many vectors, not one dense vector |
| Indexing | 9B can build an index queryable by 0.6B | Put higher compute in an offline job if query volume is high |
| Retrieval | Late interaction compares query tokens with document tokens | Use a supported MaxSim index or limit candidates before rescoring |
| API billing | No v2-late rate is published | Do not forecast managed API spend yet |
The Hugging Face late-interaction guide gives a useful scale reference from another model: a 4,874-passage example produced 608,414 token vectors and 311.5 MB of raw float32 storage, while a compressed PLAID index used 92 MB. Those numbers are not a v2-late estimate, but they show why “128 dimensions” does not equal “small index.” Token count is the important multiplier.
Indexing throughput also deserves a benchmark on your own hardware. In a real-user LocalLLaMA report, pplx-embed-v1-4b took about 45 minutes per 10,000 vectors versus 6 minutes for Qwen3-Embedding-4B on an A100 80GB. That report concerns v1, not v2-late, so it is a warning about measuring Perplexity embedding throughput, not a v2-late performance claim.
“I think it might be because pplx embed uses bidirectional attention rather than standard masked attention.” — u/Velocita84, r/LocalLLaMA
Which deployment path should you choose?
| Requirement | Best current path | Why |
|---|---|---|
| Cheap text-only RAG with a managed endpoint | Perplexity v1 API | Published prices run from $0.004 to $0.05 per 1M tokens |
| Charts, tables, scanned pages, and layout matter | Self-host pplx-embed-v2-late | The documented workflow searches rendered pages directly |
| Large corpus with frequent queries | 9B offline index plus 0.6B query encoder, or dense-first plus v2-late reranking | Separates indexing quality from query-time compute |
| Small prototype or hardware-constrained test | 0.6B checkpoint on a representative page sample | Lower active-parameter count, but still benchmark page encoding and storage |
| Managed v2-late endpoint is a hard requirement | Wait for an official API model ID and rate card | Neither is present in the current public embedding documentation |
My recommendation is to prototype the 0.6B and 9B cross-model path on 100 to 500 representative pages before building a full index. Include scanned pages, tables, multi-column layouts, and pages whose answer spans a boundary. Record recall at your target k, page-encoding throughput, raw and compressed index size, and query latency. That evidence is more useful than transferring the v1 token price to a model that is not yet sold through that API.
pplx-embed-v2-late PDF retrieval FAQ
Does pplx-embed-v2-late have an API price?
Not in the public Perplexity Embeddings API documentation checked for this guide. The published $0.004 to $0.05 per million-token prices apply to v1 standard and contextualized models.
Is pplx-embed-v2-late officially released?
Yes. Perplexity publishes 0.6B and 9B open-weight checkpoints on Hugging Face. Weight release and managed API availability are separate milestones.
Does PDF retrieval require OCR?
No, not for the visual retrieval signal. Render each page as an image and encode it as a document. OCR or extracted text remains useful for filtering, citations, accessibility, and fallback search.
Can the 0.6B model query a 9B-built index?
Perplexity’s model card says the two models share an embedding space and support that arrangement. Measure quality on your corpus because the card does not publish a cross-model retrieval delta.
Can text and image pages share one batch?
No. The model card says mixed text-and-image inputs are not supported in the same encoding batch. Keep text and image encoding calls homogeneous.
Do I still need a reranker?
Not automatically. MaxSim is already the late-interaction scoring method, but a dense first-stage retriever plus v2-late reranking can be more practical than scanning every page-token vector in a large corpus.
What is the exact storage cost per PDF page?
Perplexity does not publish a v2-late page-level storage calculator. Estimate it from the number of retained page tokens, vector precision, metadata, and index compression, then validate with a representative sample.
Should I choose 0.6B or 9B?
Use 9B when offline indexing quality is the priority and you can afford the model and indexing job. Use 0.6B for a smaller deployment or query encoder, including the documented shared-space setup against a 9B index. The benchmark gap is measurable, but the model card does not provide a universal quality or latency rule.
The go/no-go check is short: if you need a managed Perplexity endpoint with a known price today, v2-late is not ready for that requirement. If you can self-host and your PDFs contain information that OCR or text chunking loses, render a representative page set and benchmark the late-interaction pipeline before scaling it.