In an environment in which energy and infrastructure costs have become a determining factor, Gemini’s technological environment has managed to stand out with its Gemini 2.5 Pro and Flash models thanks to its vertical integration of hardware and software.
These new models not only lead in reasoning, science and multimodality benchmarks, but they do so with context windows of up to one million tokens and at generation speeds greater than 200 tokens/s.
Additionally, Google’s TPU architecture, optimized for power efficiency and performance per watt, contrasts with the industry’s near-total reliance on Nvidia GPUs, which command more than 90% of the AI accelerator market.
Let’s see in detail how its new models and its control of AI chips have placed Google in an advantageous position over its competitors.
Differences between Gemini 2.5 Pro and Flash and its direct competitors
Before diving into figures and results, it is worth distinguishing the two flavors of Gemini 2.5 and understanding its position compared to similar alternatives on the market.
Gemini 2.5 Pro is designed for demanding workloads: complex reasoning, advanced code generation, and large-scale multimodal scenarios.
Its context window of up to 128 k–1 M tokens and its design optimized for critical tasks make it a direct rival to models such as GPT-4.5 and o3-mini from OpenAI, Claude 3.7 Sonnet from Anthropic and the new Llama 3 from Meta in its most powerful version.
For its part, Gemini 2.5 Flash seeks to balance speed and cost: with minimal latencies, high throughput (>200 tokens/s) and very low prices per token, it aligns with GPT-4 Turbo, Claude 3 Haiku or Mistral Mix in its orientation towards real-time applications and high-volume services.
With this clear view of its versions and competitors, we now move on to analyze in detail the benchmarks and performance of each one.
Gemini 2.5 Pro and Flash: benchmarks and performance
Beyond the marketing speech, what really matters is how the models perform in real tests. Gemini 2.5 Pro and Flash have undergone evaluations that show their ability to compete with (and outperform) the best.
Gemini 2.5 Pro
- Reasoning and general knowledge: in the Humanity’s Last Exam benchmark, Gemini 2.5 Pro obtains 18.8% accuracy, surpassing o3-mini (14%) and Claude 3.7 (8.9%)
- Coding: achieves 63.8% in SWE-Bench Verified with custom agents, demonstrating mastery of code generation and editing tasks.
- Long Context (MRCR): scores 91.5% in co-reference tests and long dialogues of up to 128,000 tokens, significantly higher than GPT-4.5 (48.8%) and o3-mini (36.3%).
- Multimodal performance: leads in benchmarks that combine text, images, audio and video, with 81.7% in MMMU.
Gemini 2.5 Flash
- Speed: generates outputs at 248.9 tokens/s, almost double that of its predecessor Gemini 1.5 Flash.
- Controllable thinking: allows you to adjust a “thinking budget” to balance quality and latency according to the complexity of the task.
- Context window: supports up to 1 million tokens, ideal for processing long documents or long conversations without fragmenting information.
API costs and “thinking budgets”
Efficiency is not only measured in power or speed, but also in how much it costs to use it. In this section, Google has introduced a significant change that affects both large companies and small developers.
Gemini 2.5 Pro costs
$1.25 per million input tokens and $10.00 per million output tokens (sales greater than 128k) with an average “blended” price of $3.44/token when combining input and output in a 3:1 ratio.
Gemini 2.5 Flash Costs
$0.075 per million input tokens and $0.30 per million output tokens, up to 133 times cheaper than GPT-4 Turbo in inputs and 3.3 times cheaper in outputs than Claude 3 Haiku.
When used in reasoning mode on Gemini 2.5 Flash, the output ranges from $0.60 to $3.50 per million tokens, so only the “cognitive load” actually needed is billed.
What is budget thinking?
A “thinking budget” is a parameter that Google has introduced in Gemini 2.5 Flash to allow developers to control how much internal reasoning power the model uses before generating the answer.
Instead of always paying for the same level of reasoning, you can:
- Assign a token budget (between 0 and 24,576) for the “thinking” phase of the model, both via the thinking_budget parameter in the API and via the slider in Google AI Studio or Vertex AI.
- Reduce costs and latency by adjusting that budget: with a value of 0 tokens, Flash skips the complex reasoning and charges the minimum fee per output ($0.60/M tokens), achieving the lowest latency and savings of up to 600% compared to a high budget
- Pay only for what you need: when thinking is enabled, output with reasoning can cost $3.50/M tokens, but is only billed if the model actually uses those tokens to reason
TPU: Google’s strategic advantage
Google designs and manufactures its TPUs to maximize efficiency per watt, achieving up to 2.7 times better performance/Watt than the previous generation and a typical consumption of 200 W per chip in TPU v4
Google’s LCA (Life Cycle Assessment) study reveals that the latest generations of TPUs triple carbon efficiency compared to previous versions, significantly reducing CO₂ emissions from AI work
In contrast, Nvidia controls between 80% and 92% of the AI accelerator market, creating dependency on its GPUs and putting upward pressure on AI development and deployment prices.
This asymmetry of suppliers gives Google greater flexibility to adjust costs and optimize its infrastructure without depending on the limited supply of third-party chips.
Efficiency as an imperative in AI
The proliferation of large-scale models has exponentially increased data center energy consumption, resulting in high operating costs and significant environmental impact.
The need to reduce both spending and carbon footprint is therefore a priority for AI providers and users.
Sustainability specialists point out that more than 70% of the emissions associated with AI accelerator chips come from their operational electricity consumption, underscoring the relevance of improving the energy efficiency of the hardware.
The combination of cutting-edge performance and competitive pricing we see with Google Gemini 2.5 democratizes access to AI, allowing startups and SMEs to incorporate advanced solutions without incurring prohibitive costs.
This post is also available in: