Google Releases Gemini 3.7 Flash, API Prices Cut by 50%
2026-08-14 09:47
Favorite

en.Wedoany.com Reported - Google has released its new flagship model, Gemini 3.7 Flash, with coding, agentic workflows, and knowledge work as the core upgrade focus, while temporarily halving API prices. Just three weeks after the release of Gemini 3.6 Flash, this unusually short product cycle was attributed by Google to developer feedback and algorithmic improvements. For enterprise developers, the more notable aspect is the combination of enhanced intelligence and reduced inference costs: until December 31, 2026, Gemini 3.7 Flash input tokens are priced at $0.75 per million, and output tokens at $3.75 per million.

Gemini 3.7 Flash camera illustration

Starting January 1, 2027, prices will rise to $1.50 per million input tokens and $7.50 per million output tokens; context caching costs will also increase from the launch price of $0.075 per million tokens to $0.15. This means the discount is only temporary, but it gives teams deploying large-scale coding and business agents a few months to evaluate whether what Google calls "reduced retries and human oversight" translates into lower total operating costs.

Google Gemini 3.7 Flash pricing chart

Google calls Gemini 3.7 Flash "the most intelligent flagship model for coding and agents to date," emphasizing that the model is better at adapting when encountering obstacles, clarifies intent when necessary, and follows instructions with higher fidelity. In enterprise coding agent scenarios, a model that reduces unnecessary changes, recovers from errors, and executes multi-step plans more reliably can directly lower the frequency of human intervention; business agents running across documents and applications benefit similarly. Compared to Gemini 3.6 Flash's design approach of reducing inference steps, conversation turns, and tool calls, 3.7 Flash invests more effort in multi-step planning and tool invocation, aiming for more rigorous execution and fewer retries. Google DeepMind stated that the model has improved in debugging and problem-solving, can generate more practical web layouts and applications with fewer prompts, and improves reasoning and accuracy in real business workflows.

Coding benchmarks show significant improvements. On FrontierCode 1.1 Main, which measures production code quality, Gemini 3.7 Flash scored 43.6%, up from Gemini 3.6 Flash's 34.4%, and slightly higher than Google-reported figures for Claude Sonnet 5 (42.7%) and GPT-5.6 Terra (41.3%). On the long-horizon software engineering benchmark DeepSWE v1.1, 3.7 Flash reached 65.3%, compared to its predecessor's 49.0%, while GPT-5.6 Terra led with 69.6%. In web development, Gemini 3.7 Flash achieved an Elo score of 1588 on Code Arena, higher than 3.6 Flash's 1538, Claude Sonnet 5's 1541, and GPT-5.6 Terra's 1523; Google says the new model generates more practical layouts and fully functional applications with fewer prompts, and adheres more closely to reference screenshots, images, and design systems.

Broader benchmark results paint a more complex picture, which offers practical reference value for enterprises selecting models by specific workload. Gemini 3.7 Flash scored 85.8% on Terminal-bench 2.1, below GPT-5.6 Terra's 87.4%; Terra also maintains its lead on Terminal-bench 3.0 and OSWorld-2.0. Claude Sonnet 5 leads on Agent's Last Exam's multimodal desktop and operating system tasks with a 33.3% pass rate, compared to Gemini 3.7 Flash's 26.3%. In other words, Google's own data does not show 3.7 Flash fully replacing higher-priced competitors, but rather indicates the model is more competitive on coding and agent workloads while sitting at a lower price tier.

Enterprise workflow performance is a more practical test dimension. On AutomationBench, which measures enterprise workflow automation, Gemini 3.7 Flash scored 30.4%, significantly higher than 3.6 Flash's 17.0%; in Google's table, Claude Sonnet 5 scored 10.7% and GPT-5.6 Terra 23.6%. On GDP.PDF, which evaluates complex PDF understanding, the model reached 34.0%, above 3.6 Flash's 22.0%, Claude Sonnet 5's 28.0%, and GPT-5.6 Terra's 24.7%. This combination of capabilities matters greatly for enterprise agents, as real-world deployments often involve a full chain of interpreting long reports, identifying relevant information, invoking tools, updating systems, and generating documents for human review—where chain reliability is more critical than isolated reasoning benchmark scores.

HPniMGObsAAuBbD

Google is putting this approach into practice through Gemini Spark. Google AI Pro and Ultra subscribers can use 3.7 Flash in the personal AI agent Spark, improving knowledge work and tool usage across Google Workspace applications, including workflows such as consolidating files, drafting emails, and updating status documents. Enterprise users can also access the model through the Gemini Enterprise Agent Platform and Gemini Enterprise.

The launch pricing is seen as a significant attempt to embed into enterprise workflows. For comparison, Gemini 3.6 Flash's standard API pricing is $1.50 per million input tokens and $7.50 per million output tokens; Claude Sonnet 5 is $2 and $10 respectively, and GPT-5.6 Terra is $2 and $12. For autonomous agents, a single user request can trigger a long chain of model calls, reasoning tokens, and tool interactions, so a model with lower per-token costs but significantly more retries is not necessarily cheaper. Google combines lower launch token pricing with what it calls improved first-pass accuracy; if these advantages carry over to production environments, they could substantially change the operating costs of large-scale coding or document-processing agents. The metric enterprise teams should ultimately focus on is not isolated per-million-token prices, but the cost per successfully completed task.

Gemini 3.7 Flash arrives amid external discussions about whether Google is losing its edge at the AI frontier. Google has not yet released Gemini 3.5 Pro—a model that was described in May as launching the following month, then revised in July to still being in partner testing, with Thursday's announcement providing no further timeline—so the latest general-purpose Pro model from Google remains Gemini 3.1 Pro, released in February. That model was previously reported to have missed its original schedule due to failing to meet internal targets, particularly in coding, and Google has begun training what it calls its "most ambitious" model, Gemini 4.

This series of delays coincides with Google's announcement of major changes to its AI leadership. Demis Hassabis, co-founder of Google DeepMind and Nobel laureate, has transitioned to chairman of the division and chief scientist at Alphabet, no longer handling day-to-day management; former DeepMind CTO Koray Kavukcuoglu now runs the division as Senior Vice President, reporting directly to CEO Sundar Pichai. Kavukcuoglu oversees Gemini model development, frontier research, the Gemini app, and developer teams, consolidating the entire Gemini product chain under a more product-oriented executive. Chief Scientist Jeff Dean, Gemini co-leads Oriol Vinyals, Quoc Le, and Sanjay Ghemawat have left to found research startup Discovery Loop; previously, Gemini co-lead Noam Shazeer had moved to OpenAI, and Nobel laureate and AlphaFold scientist John Jumper joined Anthropic. Reuters previously reported that the slowed releases and coding shortcomings were attributed to internal disagreements, constrained compute allocation, and Google's bureaucracy.

External interpretations vary. SemiAnalysis believes Google is increasingly prioritizing the highly profitable business of providing cloud infrastructure to AI companies—including Gemini's competitors—over keeping its own models at the frontier; the analysis also claims Google effectively canceled 3.5 Pro, though Google has not confirmed this and continues to state the model is delayed. The Verge offers a more measured assessment: the departures and model delays are indeed serious, but Google retains enormous advantages through Search, Workspace, Android, cloud computing, custom AI chips, and consumer distribution; the Gemini app has over 950 million monthly active users, and its reach does not depend entirely on having the top-ranked model. Artificial Analysis's overall Model Intelligence Index rates Claude Opus 5 at 63, while Google reports Gemini 3.7 Flash at 56, an improvement over 3.6 Flash's 52. Early human preference results on Arena place 3.7 Flash at 9th overall, with web development ranking 8th.

Developers can access Gemini 3.7 Flash via Google AI Studio, the Gemini API in Android Studio, and Google's Antigravity environment; enterprises can deploy it through the Gemini Enterprise Agent Platform and Gemini Enterprise, and Google AI Pro or Ultra subscribers can use it via Spark in supported countries. Google has simultaneously updated safeguards covering chemical, biological, radiological, and nuclear risks, as well as cyberattack misuse.

The rapid leap from Gemini 3.6 Flash to 3.7 Flash demonstrates that algorithmic improvements can reach production products without waiting for a new flagship generation. For developers, this pace also brings operational requirements: production teams must still benchmark new versions against their own codebases, prompts, tool architectures, and failure modes before switching deployments. Google's own data shows substantial improvements in production-grade coding, web development, document understanding, and workflow automation, but it does not claim overall dominance. Through launch pricing, Google is effectively betting that developers will value a model that is "competitive enough with more expensive systems, while cheap enough to run repeatedly." Whether this advantage persists after prices return to full levels in January 2027 will depend on 3.7 Flash's reliability in real-world tasks, not leaderboard rankings.

This bulletin is compiled and reposted from information of global Internet and strategic partners, aiming to provide communication for readers. If there is any infringement or other issues, please inform us in time. We will make modifications or deletions accordingly. Unauthorized reproduction of this article is strictly prohibited. Email: news@wedoany.com