AIREITER

Qwen3.8-Omni-Flash API Pricing: What the Rate Card Shows

Last Updated: 2026-09-18 07:00:21

Qwen3.8-Omni-Flash pricing varies by input modality and output mode. The rates below use Alibaba Cloud's published Singapore/international listing, where audio input costs almost nine times as much as text input.

The rate card in one table

Alibaba Cloud's model and pricing documentation is the source of record for endpoint-specific rates. The international listing separates text, audio, image/video, and output charges rather than giving Qwen3.8-Omni-Flash one blended token price.

Billing itemSingapore/international price per 1M tokens
Text input$0.43
Audio input$3.81
Image or video input$0.78
Text output with text-only input$1.66
Text output after multimodal input$3.06

The figures come from Alibaba Cloud Model Studio's official pricing documentation. Confirm the same model ID and region in the console before budgeting because another Qwen model or deployment can have a different rate.

Qwen3.8-Omni-Flash is not Qwen3.8-Flash

Three similar labels appear in search results, but their rates and capabilities are not interchangeable. Match the complete model ID before using a number from any pricing table.

NameHow to treat itPricing consequence
Qwen3.8-Omni-FlashThe model covered hereUses separate text, audio, image/video, and output rates
Qwen3-Omni-FlashA different generationDo not reuse its rate card
Qwen3.8-FlashA separate Flash model lineIts text-token price is not the Omni price

The Qwen launch announcement identifies Qwen3.8-Omni-Flash as a native multimodal model. For purchasing and deployment decisions, the endpoint documentation remains the controlling source for limits and supported features.

Calculate a mixed-media request

A useful estimate applies a separate rate to each metered category. Suppose one Singapore request contains 100,000 text-input tokens, 20,000 audio-input tokens, and generates 10,000 output tokens after receiving multimodal input.

  1. Text input: 0.1 x $0.43 = $0.043.
  2. Audio input: 0.02 x $3.81 = $0.0762.
  3. Multimodal-request output: 0.01 x $3.06 = $0.0306.
  4. Estimated total: $0.1498, before taxes or account-specific adjustments.

Audio input is about 8.9 times the listed text-input rate ($3.81 / $0.43). A text-only estimate therefore understates a voice-heavy workload even when its output is short.

The billing traps that change the total

The Qwen3.8-Omni-Flash API price depends on more than total token count. Input type and whether the request is multimodal determine which rows of the rate card apply.

Audio needs its own estimate

Audio carries the largest listed input rate. For a meeting or call-analysis workflow, record audio tokens separately from the text prompt instead of multiplying all input by $0.43/M.

Multimodal input raises the output rate

Text output costs $1.66/M after text-only input and $3.06/M after image, audio, or video input. A budget that checks only input charges misses this 84% output-rate increase.

Region belongs in the cost model

Alibaba Cloud exposes region-specific deployment information, so a cost sheet should record the region beside the exact model ID. Treat the Singapore figures above as Singapore figures, not as a universal Qwen price.

Public list price is not the final invoice

Taxes, credits, promotions, and account terms can change the amount charged. The cited public listing supplies the base token rates; verify any discount or free quota in the account console rather than assuming it applies.

A production check takes four fields

Qwen3.8-Omni-Flash is worth evaluating when a workload needs combined text, audio, image, or video understanding. Before routing real users, pin four fields so a naming or endpoint mismatch does not invalidate the estimate:

  1. The complete model ID, including generation and punctuation.
  2. The region and endpoint attached to the API key.
  3. The context, output, rate-limit, and tool-support limits shown for that endpoint.
  4. The account's active credits, promotions, and expiry dates.

Run a representative request and use its returned usage data or console metrics to record text, audio, image/video, and output tokens separately. Applying the table to observed usage is more reliable than converting media minutes into tokens with an undocumented assumption.

Qwen3.8-Omni-Flash API FAQ

What is the Qwen3.8-Omni-Flash API price?

For Singapore, the official listing shows $0.43/M text input, $3.81/M audio input, $0.78/M image/video input, and $1.66/M or $3.06/M text output depending on whether the input is text-only or multimodal.

Is Qwen3.8-Omni-Flash the same as Qwen3-Omni-Flash or Qwen3.8-Flash?

No. They are separate model lines or generations. Match the complete model ID before applying a listed rate.

How is audio billed?

Audio input is billed by tokens at $3.81/M on the cited Singapore listing. It is not presented there as a flat per-minute charge.

Does image or video input change output pricing?

Yes. The listed output rate rises from $1.66/M for text-only input to $3.06/M when the request includes image, audio, or video.

The practical call

Use Qwen3.8-Omni-Flash when multimodal understanding justifies the audio and multimodal-output rates. Test one representative workload, inspect usage by modality, and compare the resulting cost per completed task rather than comparing text-input prices alone.

Sources checked