AIREITER

Gemini 3.7 Flash API Pricing Guide (Tested Specs)

Last Updated: 2026-08-13 18:47:16

Google shipped Gemini 3.7 Flash on August 13, 2026, pitching it as the strongest Flash-tier model yet for coding agents and multi-step tool use. The headline: a 65.3% score on DeepSWE v1.1 (up from 49.0% on 3.6 Flash) at an introductory price of $0.75/M input and $3.75/M output. That is exactly half of what 3.6 Flash charges at standard rates. The catch is that this pricing reportedly doubles to $1.50/$7.50 in January 2027, so the window to evaluate at the lower rate is roughly four months.

API Pricing: $0.75/$3.75 Introductory, Doubling After December 2026

Gemini 3.7 Flash enters at $0.75 per million input tokens and $3.75 per million output tokens, per the Google DeepMind launch announcement confirmed through community reports and indexed search results. Gemini 3.6 Flash's standard pricing, for reference, is $1.50/M input and $7.50/M output (AgentPedia). The 3.7 Flash intro rate matches 3.6 Flash's Batch/Flex tier but applies to standard throughput.

Community reports flag a critical detail in the fine print: the introductory rate runs through December 31, 2026, after which prices revert to $1.50/M input and $7.50/M output. One Reddit user noted that "the displayed pricing fine print causes the price to double after four months" (r/GeminiAI).

Pricing TierInput ($/M tokens)Output ($/M tokens)Notes
3.7 Flash Introductory (Aug–Dec 2026)$0.75$3.75Community-reported; doubles Jan 2027
3.7 Flash Standard (Jan 2027 onward)$1.50$7.50Same as 3.6 Flash standard
3.6 Flash Standard$1.50$7.50Current GA pricing
3.6 Flash Batch/Flex$0.75$3.75Lower-throughput tier
3.5 Flash-Lite Standard$0.30$2.50Cheapest Gemini Flash tier
Gemini Flash API pricing comparison chart

Output prices include thinking tokens. Context caching adds $0.15/M tokens at the standard tier, based on 3.6 Flash's published rates via AgentPedia. Google Search grounding and Maps grounding carry separate charges. The model retains the 1M-token context window and 64K max output from 3.6 Flash, so there are no context-limit changes that would break existing API calls.

How Much Better Are the Benchmarks?

Google-reported benchmarks show gains concentrated in software engineering, coding, and document processing. The numbers below are from Google's launch announcement, surfaced through community posts on August 13, 2026.

BenchmarkGemini 3.7 FlashGemini 3.6 FlashDelta
DeepSWE v1.165.3%49.0%+16.3 pts
FrontierCode 1.1 Main43.6%34.4%+9.2 pts
WebDev Arena (Elo)1,5881,538+50 Elo
GDP.pdf Document Processing34.0%22.0%+12.0 pts
AutomationBench30.4%~17%+13 pts
Long-Context Retrieval97%——
Gemini 3.7 Flash vs 3.6 Flash benchmark comparison chart

The 49% to 65.3% DeepSWE jump (+33% relative) is the standout result. DeepSWE tests repository-level software engineering: multi-file edits, test execution, and codebase navigation. However, the Reddit community is split on whether these numbers hold up in production. One user wrote: "Assume benchmaxxed until proven otherwise" (r/GeminiAI). Others noted that 3.6 Flash used high thinking in the comparison while some competing models used medium, which inflates scores. The 97% long-context retrieval result drew genuine interest for coding workflows where you need to find relevant code across a large codebase.

Cost vs Competing Coding Models

Early adopters on Reddit immediately compared 3.7 Flash's price-performance against cheaper alternatives. The picture is mixed: 3.7 Flash scores well, but it costs significantly more per task than budget coding models.

A commenter on the benchmark thread estimated that Gemini 3.7 Flash costs roughly 8x more than GPT-5.6 Luna on an Artificial Analysis cost-per-task basis, despite Luna having nearly identical benchmark scores. Another user said Gemini 3.7 Flash was formerly about 13x more expensive than DeepSeek V4 Flash, though DeepSeek's price has since increased. The same commenter put it bluntly: Gemini 3.7 Flash costs "more than 300% extra" compared to Luna for "nearly the same" results (r/GeminiAI).

The counterargument: Luna reportedly hallucinates more and DeepSeek V4 Flash's reasoning "makes many mistakes," though the evidence is anecdotal.

Where 3.7 Flash earns praise is speed. Users report most responses in 3 to 6 seconds, with the longest observed response (a creative-writing test) taking about one minute. Multiple commenters called it "blazingly fast."

Specs, Availability, and API Access

Gemini 3.7 Flash keeps the same context limits and feature set as 3.6 Flash, based on specs reported in community discussions and the 3.6 Flash documentation.

SpecValue
Context window1,048,576 tokens (1M)
Max output65,536 tokens (64K)
Thinking levelsLow, Medium, High
ModalitiesText, images, video, audio, PDF (input); text (output)
Knowledge cutoffMarch 2026 (3.6 Flash; 3.7 Flash cutoff unconfirmed)
Cached input pricing$0.15/M tokens (standard, inferred from 3.6 Flash)

Where to access:

  • Gemini API / Google AI Studio — developer access
  • Gemini app — consumer rollout
  • Gemini Spark — available to Google AI Pro and Google AI Ultra subscribers
  • Vertex AI — expected to follow 3.6 Flash's availability pattern

The expected API model ID is gemini-3.7-flash, following Google's naming convention from 3.6 Flash. Confirm the exact string in the Google AI for Developers documentation before deployment and pin it rather than relying on gemini-flash-latest aliases.

Should You Migrate From 3.6 Flash?

The decision depends on your workload type and how much the January 2027 price increase matters.

Migrate if your application runs coding agents, multi-step tool-use chains, or document-heavy processing. The DeepSWE jump from 49% to 65.3% is large enough to change how many agent loops succeed on the first try, and the halved intro pricing means you evaluate at lower cost. Teams already on 3.6 Flash can swap the model ID behind a feature flag without rewriting context-window assumptions.

Wait if your workload is high-volume classification, routing, or extraction where Flash-Lite already handles the job at $0.30/$2.50. The 3.7 Flash intro price is still 2.5x Flash-Lite's input rate, and the January doubling widens that gap.

Plan around the price change. Budget for $1.50/$7.50 starting January 2027, not the current intro rate. If your unit economics only work at $0.75/$3.75, the model becomes 2x more expensive in four months.

FAQ

Is Gemini 3.7 Flash's $0.75/$3.75 pricing permanent?

No. Community reports indicate the rate is introductory, valid through December 31, 2026. Starting January 2027, prices reportedly double to $1.50/M input and $7.50/M output, matching 3.6 Flash's standard rates.

Is Gemini 3.7 Flash available on Vertex AI?

Vertex AI availability has not been independently confirmed at launch. Gemini 3.6 Flash was available through Vertex AI at GA, so 3.7 Flash is expected to follow. Check the Google Cloud console for the current model list.

How does 3.7 Flash compare to Gemini 3.1 Pro?

Reddit users report that 3.7 Flash is competitive with or better than Pro-tier models on specific agent and coding benchmarks, particularly agentic document analysis. For pure coding-agent loops at the intro price, 3.7 Flash offers better cost-performance; for complex reasoning and planning, Pro remains the stronger choice.

Are thinking tokens billed separately?

No. Thinking tokens are billed as output tokens at the standard output rate ($3.75/M during the intro period). Using high thinking increases output costs proportionally, so use medium or low for high-volume workloads where extra reasoning is unnecessary.