AIREITER

Prime Inference Pricing and API Guide (2026)

Last Updated: 2026-10-03 07:21:29

Prime Inference is live, OpenAI-compatible, and built for sustained agent traffic—not merely a future product announcement. The catch is pricing transparency: Prime Intellect’s launch post explains the serving stack in detail, but the complete per-model rate card is currently easier to find through live pricing directories than on a conventional official pricing page.

Prime Intellect's Prime Inference launch page

Prime Inference is live, but pricing is split across sources

Prime Intellect formally released Prime Inference on October 2, 2026. The product includes serverless endpoints for variable demand and reserved capacity for sustained workloads, with the public API separated from the underlying model fleet so capacity can move or fail without changing the client endpoint.

This is a production release rather than a waitlist. Prime Intellect says its first public deployment, GLM-5.3, went live on OpenRouter on September 22, while direct clients can point OpenAI-compatible SDKs to https://api.pinference.ai/api/v1.

The pricing caveat is straightforward: the official launch page promises unified billing and team-level usage tracking but does not publish a complete rate table. Current model prices therefore need to be checked in the authenticated dashboard or models endpoint before committing spend. Public directories are useful snapshots, not contractual quotes.

What Prime Inference costs right now

Prime Inference pricing is model-specific and usage-based. Input and output tokens are billed separately; output can cost several times more than input, so a cheap-looking input rate may be misleading for coding agents or reasoning workloads that generate long answers.

The following snapshot was recorded on October 3, 2026 from endpoints.run’s Prime Intellect catalog, which listed 64 endpoints. Verify each rate in Prime Intellect before purchase.

Model IDInput / 1M tokensOutput / 1M tokensListed contextBest fit
meta-llama/Llama-3.2-1B-Instruct$0.03$0.2060KClassification and simple extraction
qwen/qwen3.7-flash$0.03$0.131MLow-cost, long-context general work
z-ai/glm-5.3-flash$0.15$0.501MFaster agent and tool workflows
qwen/qwen3-coder-next$0.30$1.50262KCoding and tool use
deepseek/deepseek-v4.1-flash$0.30$1.201MLong-context reasoning and vision
z-ai/glm-5.3$1.40$4.401MFlagship GLM agent workloads
deepseek/deepseek-v4-pro$1.91$3.831MHigher-end reasoning
moonshotai/kimi-k3$3.45$17.251MPremium long-context tasks

A second directory, Compute Prices, displayed 111 models on the same date and an exact low of $0.027 per million input tokens for Llama 3.2 1B. The catalog-count mismatch is a reason to query the live API rather than treat either directory as a fixed inventory: listings can include different namespaces, gateway models, or update at different times.

Three realistic cost examples

Token cost can be estimated with:

(input tokens ÷ 1,000,000 × input rate) + (output tokens ÷ 1,000,000 × output rate)

Monthly workloadModel and token volumeEstimated token bill
Lightweight support botQwen3.7 Flash; 50M input + 10M output$2.80
Coding agentQwen3 Coder Next; 100M input + 40M output$90.00
Output-heavy flagship agentGLM-5.3; 200M input + 100M output$720.00

These calculations cover listed token charges only. They do not include reserved-capacity contracts, sandbox execution, training, evaluations, or GPU rental. Prime Intellect sells those as separate services, so an end-to-end agent bill can contain more than inference tokens.

Long-context support also does not mean every request consumes one million tokens. Billing follows actual processed tokens. However, an agent that repeatedly resends a 140K-token history can accumulate input charges quickly unless prefix caching or application-side context management reduces repeated work.

The API setup takes one endpoint change

Prime Inference exposes an OpenAI-compatible endpoint, so an existing OpenAI SDK integration generally needs a new base URL, API key, and model identifier. Prime Intellect’s launch post gives the base URL as https://api.pinference.ai/api/v1.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_PRIME_API_KEY",
    base_url="https://api.pinference.ai/api/v1",
)

response = client.chat.completions.create(
    model="z-ai/glm-5.3",
    messages=[
        {"role": "user", "content": "Write a haiku about KV caches."}
    ],
    stream=False,
)

print(response.choices[0].message.content)

The equivalent cURL request is:

curl https://api.pinference.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $PRIME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "z-ai/glm-5.3",
    "messages": [
      {"role": "user", "content": "Write a haiku about KV caches."}
    ]
  }'

Prime Intellect also documents a CLI path: install the prime tool, authenticate with prime login, then run prime inference chat. Before deploying, retrieve the current model list and copy the exact model ID; directory prefixes vary across catalogs.

What the launch architecture changes for agent workloads

Prime Inference is most differentiated when prompts are long, sessions return repeatedly, and tool calls must remain valid. Prime Intellect says a typical internal agent turn adds about 6K tokens to a 140K-token prompt, which makes cache retention and routing as important as raw decode speed.

The official launch report supplies several concrete measurements from its GLM-5.3 deployment on GB200 NVL72 hardware:

  • Separating prefill and decode workers reduced p90 inter-token latency by nearly 40% in Prime Intellect’s tests.
  • At a target of 100 end-to-end tokens per second per user, a tested 1:4 prefill/decode topology served 66 sessions per prefill group at 101 tokens per second per user.
  • Reducing the prefill budget from 8K to 4K tokens per GPU cut median queue wait from 550 ms to 110 ms and reduced median time to first token by roughly 20%.
  • NVFP4 compression increased cache capacity from 1.09 million to 1.63 million cached tokens per decoder—about 50%—at the same memory budget.
  • A block-major cache layout reduced transfer descriptors by about 10× and mean transfer time from 146 ms to 78 ms in a separate TP8 comparison.

Those figures are workload- and configuration-specific, not universal latency guarantees. Their practical value is that Prime Intellect publishes the bottlenecks it optimized: returning-session cache reuse, admission delay, prefill/decode interference, and cache transfer overhead.

Tool use deserves a separate test. Prime Intellect found cases where undeclared tools were silently discarded or arguments had incorrect types, then added constrained decoding and schema tests to its GLM serving path. The company reports a near-zero tool-call error rate, but buyers should still replay their own nested schemas, nullable arguments, streaming calls, and malformed requests before switching production traffic.

Serverless or reserved capacity: choose by traffic shape

Serverless is the sensible starting point for most teams because it converts uncertain demand into token charges. Reserved capacity becomes relevant when sustained concurrency, predictable latency, or a custom deployment matters more than avoiding idle capacity.

WorkloadBetter starting optionReason
Prototype or intermittent chatbotServerlessNo need to reserve idle GPUs
Bursty evaluation jobsServerless nowBatch and async lower-price inference is listed on the roadmap, not as a current launch feature
Steady high-concurrency agent serviceReserved quoteCapacity planning may matter more than nominal token rates
Fine-tuned or LoRA modelConfirm availability firstOne-click dedicated deployment of fine-tuned models is described as roadmap work in the launch post
Strict contractual SLA or region requirementSales reviewPublic launch material does not state a universal contract, region matrix, or standard reserved price

Do not assume “reserved” automatically costs less. Compare the quoted capacity cost with measured serverless spend at representative concurrency, including idle periods and peak headroom.

The facts Prime Intellect still does not publish clearly

Prime Inference has unusually detailed infrastructure disclosure, but several purchasing facts remain unclear in public, indexed material:

  • a complete official per-model rate card;
  • cached-input and reasoning-token billing rules for every model;
  • public rate and concurrency limits;
  • a region-by-region processing map;
  • standard data-retention terms for inference requests;
  • a universal formal SLA and remedies;
  • minimum terms and prices for reserved capacity;
  • which catalog entries are directly Prime-hosted versus routed gateway models.

Unknown does not mean unsupported. It means these items should be confirmed in the dashboard, documentation, model endpoint, contract, or sales response rather than inferred from OpenAI compatibility.

Prime Inference FAQ

Is Prime Inference officially available?

Yes. Prime Intellect announced the public release on October 2, 2026 and published an API base URL, CLI command, and serverless/reserved product split.

Does Prime Inference have a free tier?

The public launch post does not specify free Prime Inference credits. Prime Intellect has free or zero-priced options elsewhere in its training stack, but those should not be treated as an inference free tier.

Is Prime Inference compatible with the OpenAI SDK?

Yes. Set the OpenAI-compatible base URL to https://api.pinference.ai/api/v1, provide a Prime API key, and use a currently available model ID.

Does Prime Inference support tool calling?

The GLM-5.3 serving path supports tool calling, and Prime Intellect describes grammar enforcement and regression tests for argument schemas. Support should still be checked per model because model capabilities differ.

How many models are available?

Public directories displayed 64 and 111 entries on October 3, 2026. Because the counts differ, the live Prime model endpoint or dashboard is the reliable source for the inventory available to an account.

Is Prime Inference cheaper than renting GPUs?

Serverless is usually easier to justify for variable traffic; rented or reserved GPUs can become economical at high, stable utilization. The answer depends on measured tokens, concurrency, latency targets, and idle capacity—not token price alone.

A five-minute go/no-go check

Prime Inference is worth testing for teams that need open-model access, long-context agent serving, or a path from Prime Intellect’s training stack into production. It is not ready for blind procurement from a public rate card alone.

  1. Query the live model list and record the exact input, output, cache, and reasoning rates.
  2. Replay a representative long conversation and measure first-token latency, generation speed, and total billed tokens.
  3. Run the application’s real tool schemas in streaming and non-streaming modes.
  4. Ask for retention, region, rate-limit, SLA, and reserved-capacity terms if the workload is sensitive.
  5. Compare measured serverless spend with a reserved quote only after observing normal and peak traffic.

The deciding trade-off is clear: Prime Inference provides credible serving engineering and a low-friction API, while price and contract verification still require one more step than the launch page itself provides.