Google Cuts API Costs in Half With Launch of Gemini 3.7 Flash
Google slashes token pricing and boosts benchmark scores across software engineering and document processing
Google has cut its API pricing in half while launching Gemini 3.7 Flash, replacing its 3.6 build after only a few weeks with a new model optimized for software engineering and autonomous AI agents.
Company executives cited direct feedback from developers using prior iterations as the primary driver behind the update. The architecture introduces targeted modifications to core algorithms and system workflows, prioritizing software engineering tasks, web development, and agentic workflows.
Along with speed and reasoning gains, the release significantly lowers operating expenses for engineering teams. Gemini 3.7 Flash costs half as much per million tokens as the previous generation, reducing the commercial hurdle for developers integrating AI into production software.
Internal evaluation metrics indicate marked improvements over Gemini 3.6 Flash in code debugging and troubleshooting tasks. On synthetic programming evaluations, the updated build achieved a 43.6% score on FrontierCode 1.1 Main compared to 34.4% for the prior generation, while rising to 65.3% on DeepSWE v1.1 from 49.0%.
Synthetic benchmarks such as DeepSWE v1.1 measure how effectively an artificial intelligence model can autonomously resolve complex software engineering problems, including bug fixing and repository-level code modifications. Higher scores indicate that an automated system can handle multi-step programming challenges with less human supervision.
Gemini 3.7 Flash Outperforms Predecessor Across Industry Benchmarks
Performance updates extend beyond software engineering into specialized domain processing. Google reported that Gemini 3.7 Flash registered a 34.0% score on the GDP.pdf document evaluation test compared to 22.0% for version 3.6, reflecting improved analysis of complex materials in finance, law, and biosciences. On AutomationBench, performance increased to 30.4% from 17.0%.
The model launched immediately with promotional pricing tiers active through the end of the year. Developers pay $0.75 per million input tokens and $3.75 per million output tokens during the initial phase. Rates transition on January 1, 2027, to $1.50 per million input tokens and $7.50 per million output tokens.
Token-based pricing schedules dictate the computational cost of running continuous AI agents that continuously read codebases and write multi-line responses. Promotional rates allow software enterprises to test large-scale agent deployments before standard production pricing takes effect.
The pricing schedule allows software developers to implement the model immediately under discounted terms through December before standard rates take effect in 2027.








