AIREITER

Liquid AI d1 API Review: Open Models vs Hosted Pricing

Last Updated: 2026-10-07 19:16:46

The hosted API remains the simplest route at $0.04 per million input tokens, but the new open-weight checkpoints make local inference practical, with a sharp quality gap between the 3B and 600M models.

The three Liquid d1 options are not interchangeable

Liquid AI now offers three relevant choices: the proprietary hosted d1 service, the open-weight d1-3B, and the experimental open-weight d1-omni-600M. Liquid AI has not stated that its hosted model is identical to either downloadable checkpoint, so benchmark or latency figures must stay attached to the model that produced them.

ChoiceAccess and priceInputsPublished contextBest fit
Hosted d1Liquid API; $0.04 per 1M input tokens, no output-token chargeText and images direct from Liquid66K listed by VercelFast integration, negligible inference bill, no model operations
d1-3BDownloadable weights; no per-token model feeText, JSON, and images32,768 tokensBest local quality, vision, production evaluation
d1-omni-600MDownloadable experimental weights; no per-token model feeText plus image, or text plus audio16,384 tokensSmall footprint, voice-command experiments, constrained edge devices

All three answer typed questions rather than write prose. noul returns a yes probability, choice returns probabilities over named options, and score returns a probability-weighted position on an ordered rubric. These models suit routing, moderation, inspection, guardrails, and scoring; they do not replace a chat or coding model.

The recommendation is direct: use hosted d1 for a pilot, choose d1-3B when data must remain local or latency must avoid a network hop, and treat d1-omni-600M as an experiment unless audio or footprint is the decisive constraint.

API pricing versus local cost

The Liquid AI d1 API price is $0.04 per million input tokens, with no output-token charge because the service returns decisions rather than generated text. A workload of 100 million input tokens costs $4 before any gateway fees; one billion input tokens costs $40.

At $0.04/M tokens, self-hosting rarely wins on inference cost alone. Choose local deployment for privacy, offline use, on-device latency, customization, or sustained utilization.

The two open checkpoints have no official hosted per-token price. Their Hugging Face pages showed no inference provider at review time, so the hosted d1 price must not be presented as d1-3B or d1-omni-600M pricing.

The hosted service also bills images as input. Liquid AI specifies 1.5 tokens per 32×32-pixel patch, making a 1024×1024 image 1,536 tokens before question text. At $0.04/M, the image portion costs about $0.00006144; each question is billed as its own prompt and includes its associated image.

Open-weight does not mean unrestricted or cost-free

Both repositories use Liquid AI's lfm1.0 license, not Apache 2.0 or MIT. Liquid AI's general pricing terms say its downloadable models are free for commercial use below $10 million in annual company revenue; larger organizations should verify the current commercial terms rather than infer permission from the phrase “open-weight.”

For local budgeting, include GPU or device acquisition, idle capacity, monitoring, updates, and staff time. The hosted bill is variable and tiny; local cost is mostly fixed.

d1-3B is the quality choice; omni is the footprint choice

The official Open d1 announcement reports a Decision Index v0.2.1 score of 48.57 for d1-3B and 15.95 for d1-omni-600M. Liquid AI ran both through the official scorer, but the results were not official leaderboard submissions.

Bar chart comparing Decision Index scores for Liquid AI d1-3B and d1-omni-600M

The smaller model is not uniformly poor. Across seven public text benchmarks, Liquid AI reports means of 82.9 for d1-3B and 78.4 for d1-omni-600M. Omni led on Civil Comments (95.8 versus 93.0) and PAWS-X (79.5 versus 76.9), while 3B led on the other five listed tasks. The much wider Decision Index gap indicates that a few strong classification slices do not make the 600M checkpoint a general substitute for 3B.

Specificationd1-3Bd1-omni-600M
Detailed parameter count3.12B587M
Full-precision repository/weight footprint6.27 GB repository; 6.25 GB weightsF32 model, about 0.6B parameters
Context32,76816,384 shared across modalities
Image text allowanceWithin model contextState and question text cut to 896 tokens with images
AudioNoOne 16 kHz mono clip, up to 30 seconds
Image + audio togetherNot applicableNot supported; raises ValueError
Official speed tableYesNo; early research release

The d1-omni-600M model card says the model was trained in float32. Float16 preserved the top answer on 243 text, 214 image, and 416 audio rows, while bfloat16 changed the top answer on 0.8% of text rows and 1.7% of audio rows. That makes dtype a validation variable, not merely an optimization switch.

Liquid AI d1-3B model page showing the official open-weight repository

Local deployment has a short setup and a long validation tail

Both model cards expose system_one for one state and system_one_batch for packed requests. The shortest d1-3B path uses Transformers 5.14 or later, PyTorch, TorchVision, Pillow, and repository-supplied custom code.

  1. Install the documented dependencies:
pip install "transformers>=5.14" torch torchvision pillow
  1. Load the model with remote repository code enabled:
import torch
from transformers import AutoModel

model_id = "LiquidAI/d1-3B"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32

model = AutoModel.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=dtype,
).to(device)
  1. Submit named decisions rather than a chat prompt:
questions = {
    "queue": {
        "type": "choice",
        "instructions": "Which team should handle this ticket?",
        "criteria": {
            "billing": "Charges, invoices, and refunds",
            "technical": "Application or website faults",
            "fraud": "Suspected unauthorized use",
        },
    }
}

result = model.system_one(
    "I was charged twice this month; refund one charge.",
    questions,
)
print(result["answers"]["queue"])

trust_remote_code=True means the repository's Python executes in the serving process. Pin a reviewed commit before production rather than loading a moving branch. The d1-3B repository also links quantizations for llama.cpp, Ollama, and LM Studio-compatible runtimes; a quantized community artifact is a separate supply-chain and accuracy decision from the official BF16 weights.

The d1-3B model card reports warm-call times of 8 ms for one question on an RTX 4090, 30 ms on an Apple M5 Pro, and 50 ms on a Jetson Orin Nano. A 3.4K-token state took 102 ms, 640 ms, and 1,640 ms on those same devices. The GPU figures are medians over 20 runs; the 8 ms 4090 result used compiled CUDA graphs, and a new input shape pays compilation or kernel-selection overhead.

No official d1-omni-600M speed table exists. A browser port reported roughly 180 ms per comment for four decisions:

“I didn't test it against other machines yet, just my MBP M4 Pro.” — u/FinancialAd1961 on Reddit

The single-device result demonstrates browser feasibility, not performance across browsers, GPUs, quantizations, or batch sizes.

API access is easier, but model identity still matters

Direct hosted access uses model d1 at https://api.liquid.ai/decisions/v1/systemone. Vercel uses liquid/d1; its page lists a 66K context and the same $0.04/M input rate. Liquid's October 5 announcement said Vercel and OpenRouter were text-only at that point, while the direct Liquid API accepted images.

The open checkpoints use local Python methods with the same decision concepts, but interface similarity does not establish behavioral equivalence. Their probabilities can differ enough to move a case across an automation threshold. Record the exact model ID, revision, dtype, probability, and threshold for every evaluated branch.

A migration should therefore follow four steps:

  1. Build a labeled set from the actual production distribution, including ambiguous and costly errors.
  2. Run the same states, question schema, and criteria through each candidate.
  3. Select thresholds from false-positive and false-negative costs, not from a global benchmark score.
  4. Shadow the winner before allowing destructive, financial, access-control, or safety actions.

Which Liquid d1 should you choose?

The hosted API is the default recommendation for most teams. Its token price is too low for a small pilot to justify a local serving project, and it avoids driver, quantization, warm-up, and capacity work.

RequirementPickReason
Fastest path to productionHosted d1Managed endpoint and explicit input pricing
Data must not leave the device or networkd1-3BStrongest open checkpoint and local execution
Best published open-model decision qualityd1-3B48.57 Decision Index versus 15.95
Audio command classificationd1-omni-600MOnly option here with audio, capped at 30 seconds
Smallest experimental footprintd1-omni-600M587M parameters, but no official latency table
Open-ended explanations or generated actionsNone of themAdd a generative model after the decision stage

Choose d1-3B over omni unless the 600M footprint or audio path is indispensable. A fivefold parameter reduction is attractive, but it does not erase the 32.62-point Decision Index gap or omni's experimental status.

Liquid AI d1 FAQ

Is Liquid AI d1 free?

The hosted d1 service is priced at $0.04 per million input tokens with no output-token charge. The open checkpoints can be downloaded without a per-token fee, but the LFM 1.0 commercial terms and local compute costs still apply.

Are d1-3B and d1-omni-600M the hosted d1 API model?

Liquid AI has not documented either checkpoint as identical to hosted d1. Treat them as related products with separate model IDs, contexts, benchmarks, and deployment paths.

Can I run Liquid d1 with Ollama?

The d1-3B model page links quantizations intended for llama.cpp, Ollama, LM Studio, and compatible applications. Verify the quantizer, source revision, decision API support, and accuracy before replacing the official Transformers path.

Does d1-3B support audio?

No. d1-3B accepts text, JSON, and images. d1-omni-600M accepts text plus images or text plus one audio clip of up to 30 seconds, but it cannot accept images and audio in the same request.

How much VRAM does local d1 need?

Liquid publishes a 6.25 GB BF16 weight file for d1-3B but does not give one universal VRAM requirement. Runtime overhead, image encoder state, context length, batch size, precision, and quantization all change the total; measure the intended configuration rather than equating file size with peak VRAM.

Whichever route you choose, validate probability thresholds on labeled production data before automating consequential decisions.

Related reading: the hosted Liquid AI d1 API review covers the direct endpoint, image billing, and typed request contract in more detail.