Skip to content
Developer312
AI News4 min read

Google ships Gemini 3.7 Flash at $0.75/$3.75 — half the price, three weeks after 3.6

Google launched Gemini 3.7 Flash today at $0.75/M input and $3.75/M output — half what 3.6 Flash cost three weeks ago. 1M context, coding and agent focused, locked at this rate through year-end. The cheap-fast tier price war just got real.

By Developer312Published August 13, 2026Report an error

TL;DR: Google launched Gemini 3.7 Flash today — three weeks after 3.6 Flash — at $0.75/M input and $3.75/M output, a 50% cut from 3.6 Flash's launch price. The model targets coding and agent workloads with a 1M context window. It is not positioned as a frontier model; it is positioned as the cheapest viable workhorse that still beats most of the field on tool-use and code generation.

Key Takeaways

  • Gemini 3.7 Flash launched three weeks after 3.6 Flash at $0.75/M input and $3.75/M output — half of 3.6's launch price, locked through year-end 2026.
  • The model targets coding and agent workloads with a 1M-token context window. Google positions it as 'most intelligent workhorse,' not as a frontier model.
  • Benchmark coverage (OfficeChai) places it below Claude Opus 5 and GPT-5.6 Sol on reasoning, but competitive with GPT-5.6 Luna and Claude Haiku 4.5 on the cheap-fast tier.
  • Google has reportedly shifted internal priority away from Gemini 3.5 Pro toward Gemini 4 — Pro-line investment at this generation is winding down.

What shipped

Gemini 3.7 Flash is the third Google model release in the Flash line this quarter, and the second in three weeks. Google is explicitly framing it as a response to developer feedback on 3.6 Flash, which shipped in late July. The official Google blog describes 3.7 Flash as "our most intelligent workhorse model yet for coding and agents" — note the phrasing. Not frontier. Workhorse.

The model has a 1M-token context window and is available through the Gemini API, AI Studio, and Vertex AI. Google's pricing is $0.75 per million input tokens and $3.75 per million output tokens, locked through the end of 2026. That is half what 3.6 Flash launched at.

The timing of the cut matters more than the headline number. OpenAI cut GPT-5.6 Luna 80% at the end of July. Anthropic has been moving Sonnet pricing down through successive versions. Alibaba launched Qwen3.8-Max in August at $2/$6. The cheap-fast tier is now in a price war, and Google is choosing to compete on rate cards rather than capability claims for the Flash line specifically.

Where it sits in the field

OfficeChai's benchmark analysis puts Gemini 3.7 Flash "not quite at the frontier, but good for some use cases" — specifically coding, multi-step tool use, and high-volume summarization. It is not going to win head-to-heads against Claude Opus 5 or GPT-5.6 Sol on reasoning benchmarks. That is by design.

The relevant comparison is against the cheap-fast tier: GPT-5.6 Luna, Claude Haiku 4.5, Mistral Small 4, Llama 4 Maverick, and the small Mistral/DeepSeek open weights. On cost-per-useful-output, Gemini 3.7 Flash at $0.75/$3.75 is now among the cheapest serious options for a high-volume coding or agent workload. The catch is that Google's lock-in through year-end means the price is temporary — if you build cost models around it, plan for a 2027 reset.

What builders should take from this

If you are running a coding agent in production today: The price cut is real, not vapor. For high-volume code generation, summarization, and tool-use workloads that do not need frontier reasoning, Gemini 3.7 Flash is now worth re-benchmarking. The lock-in through 2026 means your cost model can assume the rate for the rest of the year.

If you are building a multi-model architecture: Google's pricing is now aggressive enough to consider Gemini 3.7 Flash as a router target for cheap-fast requests. Nvidia's Switchyard play (released earlier this week) is built around this exact idea — route cheap-and-fast requests to a model like 3.7 Flash, expensive-and-correct requests to a frontier model.

If you were considering Gemini 3.5 Pro: Don't. Google has reportedly shifted internal priorities away from 3.5 Pro development toward Gemini 4. If you were betting on continued investment in the Pro line at this generation, the signal is that Google is consolidating fast-tier investment and pushing Pro-tier to the next version.

The take

The cheap-fast tier is the actual battleground for AI infrastructure revenue, and Google just made the most aggressive move of the cycle. The $0.75/$3.75 pricing is below Anthropic Haiku, below Mistral Small, and competitive with open-weight self-hosted options once you account for ops cost. Three weeks between 3.6 Flash and 3.7 Flash is also the new tempo for model iteration in the cheap tier — expect this to be the new floor, not the new ceiling.

Sources

  1. [1]Gemini 3.7 Flash: our most intelligent workhorse modelGoogle (2026-08-13)
  2. [2]Google launches Gemini 3.7 Flash AI model for codingQuartz (2026-08-13)
  3. [3]Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at...MarkTechPost (2026-08-13)
  4. [4]Gemini 3.7 Flash benchmarksOfficeChai (2026-08-13)
  5. [5]Google Drops Gemini 3.7 Flash Pricing to $0.75 Per Mtok in Pivot to Gemini 4Hugging News (2026-08-13)

Get the next briefing

Signal-first AI briefings, weekday mornings.

One concise briefing with three signals, why they matter, and one action to take.

Free. No spam. Unsubscribe anytime. · Weekday mornings.

Share this article

Related Articles