AIREITER

Gemini Omni Flash Pricing & API Access Guide

Last Updated: 2026-07-01 07:24:23

Gemini Omni Flash's real API price is about $0.10 per second of 720p video — that's $17.50 per million output tokens, billed at 5,792 tokens per second, per Google's own Gemini API pricing page, checked July 2026. There's no free tier for this model. The model ID — easy to miss since it's buried in the docs — is gemini-omni-flash-preview, still tagged "preview," so pricing and access details can shift before general availability. It ships through two separate doors, Google AI Studio and Vertex AI, that don't share code. If you've seen a "$0.20–0.60 per second" estimate somewhere, that was a pre-launch guess written before Google published real numbers; it's outdated now.

What Gemini Omni Flash Actually Is

Per Google's launch post, Gemini Omni Flash is the first model in a new "Omni" family — an any-to-any transformer taking text, images, audio, and video as input and returning video grounded in Gemini's world knowledge. It's free for Google AI Plus, Pro, and Ultra subscribers through the Gemini app, Google Flow, and YouTube Shorts/Create; developer API access followed in late June 2026, per VentureBeat.

Its standout feature is conversational, multi-turn editing: describe a change, it applies the edit while keeping characters and scenes consistent, and you keep iterating in the same thread.

The DeepMind model card lists the safety guardrails — SynthID watermarking and C2PA content credentials on every clip, no lip-syncing a still photo of a real person to audio, restricted voice-editing — and three known weak points: consistency across multi-step edits, complex-motion scenes, and accurate on-screen text. We tested all three below.

We Tested the Multi-Turn Editing Ourselves

We ran it through Google Flow with a two-round edit sequence: generate a woman holding a handwritten cardboard sign, then edit that same clip to switch it to nighttime with a neon sign and have her turn and walk toward camera. Both output files came back at 720p, 24fps, 8 seconds, with native stereo audio — matching what third-party integrators have reported, now independently confirmed on our own files.

The on-screen text on the cardboard sign — "OPEN TODAY" — rendered perfectly, no garbling, for the full 8 seconds. That's notable because accurate text rendering is one of the weaknesses the model card itself admits to.

The character consistency claim holds up well in practice: same face, same red puffer jacket, same white sweater, same hair, carried over cleanly into a completely different lighting setup with a new neon sign added to the scene. Two things didn't go as instructed, though:

  • The background store signage broke text rendering — the "COFFEE" sign in the window came back reading "COFFFEE" (extra letter), even though the foreground cardboard sign in round one was flawless. Text accuracy seems to depend heavily on how prominent and how central the text is to the frame, not just whether the model can render text at all.

  • The requested "turn and walk toward camera" motion barely happened. Watch the clip above — what we got instead was a subtle camera-distance shift with the subject staying mostly stationary, not the clear turn-and-walk called for in the prompt. This lines up with the "complex-motion scenes" limitation Google discloses on the model card; multi-step body motion doesn't yet track the instruction as reliably as a static-to-static lighting or scene change does.

One more thing worth budgeting for: a single mistyped edit consumed one of our available generations on Flow, which is enough to burn through a testing session faster than expected. Google doesn't publish an exact per-plan generation quota for Omni Flash, so if you're testing seriously, expect to spend more attempts than you plan for and don't assume you can retry a botched prompt for free.

Gemini Omni Flash Pricing, Broken Down

Here's the rate card, per Google's Gemini API pricing page — no free tier, Standard rate only:

Price

Input (text / image / video / audio)

$1.50 per 1M tokens

Output — text

$9.00 per 1M tokens

Output — video

$17.50 per 1M tokens

Video isn't billed by the second directly — it's billed by output tokens, and Google specifies the conversion: 5,792 tokens per second of 720p video. Do the math ($17.50 ÷ 1,000,000 × 5,792) and you land on roughly $0.101 per second, which Google's docs round to "approximately $0.10 per second."

What that means for a real clip, at 720p:

Clip length

Approx. cost

4s

$0.41

5s

$0.51

6s

$0.61

8s

$0.81

10s

$1.01

Add input tokens on top — a short text prompt plus one or two reference images typically adds a few cents at $1.50/1M — and an 8-second clip lands close to $0.85 total. Google's page only documents the token-to-second conversion for 720p; 1080p and 4K aren't listed separately, so budget those as likely more expensive and re-check the live page before a production run.

How to Actually Get API Access

Only two channels are actually confirmed official — be skeptical of any third-party guide claiming broader resale access before Google's own channels are fully open:

  1. Google AI Studio — get a key at aistudio.google.com/apikey, export it as GEMINI_API_KEY, and call the model as gemini-omni-flash-preview through the Interactions API (pip install -U google-genai or npm install @google/genai). Fastest path for testing.

  2. Vertex AI — same underlying model, fronted by a separate enterprise-grade endpoint that additionally requires a GCP project ID, a region/location setting, and IAM-based auth instead of a plain API key. Budget 30-60 minutes of setup if you haven't touched Google Cloud billing before.

Worth planning around: AI Studio and Vertex AI are separate Google product surfaces — different docs, different SDKs, and API-key auth versus GCP IAM auth respectively. They aren't interchangeable by design, so pick whichever platform you'll actually deploy on before you build, rather than prototyping against AI Studio and assuming a copy-paste move to Vertex later.

One more thing worth knowing before you start: per Google's Interactions API get-started guide, video generation defaults to background=True for long-running tasks rather than a synchronous response, which also means it isn't OpenAI chat-completions compatible. A minimal Python pattern — illustrative; check the current google-genai SDK reference for exact field names before shipping it:

from google import genai
import time

client = genai.Client()  # reads GEMINI_API_KEY from env

interaction = client.interactions.create(
    model="gemini-omni-flash-preview",
    input="A drone shot pulling back from a lighthouse at sunset",
    background=True,
)

while interaction.status not in ("completed", "failed"):
    time.sleep(3)
    interaction = client.interactions.get(interaction.id)

if interaction.status == "completed":
    with open("output.mp4", "wb") as f:
        f.write(interaction.output_video.read())

There's no drop-in chat-completions shape here — if your existing integration code assumes OpenAI's synchronous response format, you're rewriting the request/response handling, not swapping a model string.

Gemini Omni Flash vs. Single-Shot Video Models

Pick Omni Flash for the conversational edit loop. For a straight text/image-to-video job at a fixed resolution and duration, a single-shot model is simpler and cheaper to reason about:

Best for

Access

Gemini Omni Flash

Multi-turn conversational editing, character consistency across edits

Google AI Studio or Vertex AI directly

Veo 3.1 / Sora 2 / Seedance 2.0

One-shot text/image-to-video at a fixed spec

Direct from each vendor, or bundled on relay platforms such as AIReiter

Omni Flash isn't on AIReiter or similar bundled relay platforms yet — for its multi-turn workflow, going straight to Google's own API remains the simplest option.

FAQ

What is the Gemini Omni Flash model ID?

gemini-omni-flash-preview. Use this exact string when calling it through the Interactions API in Google AI Studio or Vertex AI. It's still a preview model, so the ID could change at general availability.

How much does Gemini Omni Flash cost per video?

About $0.10/sec at 720p, so a typical 4-10 second clip runs $0.50-$1.

Is there a free tier for Gemini Omni Flash?

No. Google's pricing page lists "Not available" for both input and output on the free tier — Standard (paid) tier only.

Can I use the same code on AI Studio and Vertex AI?

Not directly. They're separate Google product surfaces with different SDKs and auth (API key vs. GCP IAM), so code written for one needs real rework to move to the other. Pick your deployment target before you build.

Does Gemini Omni Flash work with OpenAI-style chat completions code?

No. Video generation runs through async task creation and polling (background=True plus interactions.get()), not a synchronous chat-completions response format.

Does the multi-turn character consistency actually work?

In our own two-round test through Google Flow, yes for appearance — the same face, jacket, and hair carried cleanly through a full day-to-night scene change. It struggled more with following complex motion instructions (a requested turn-and-walk barely registered) and with rendering background text accurately, both of which match limitations Google discloses on the model card.