🚧 Beta — whatiswhat.ai is actively being built out. You may notice gaps or rough edges — thanks for your patience!

News

Google DeepMind's New Reasoning Model Outperforms GPT-5 on Math

May 3, 2026, 11:15 AM ET 3 min read By Kyle Van Buren
AI-assisted, human-reviewed: Portions of this article were drafted with AI tools for research efficiency. Every claim was reviewed and edited by Kyle Van Buren, Founder of whatiswhat.ai. Learn about our process.
TL;DR
  • Google DeepMind released Gemini Reasoning 1.0, a new model focused on mathematical and scientific reasoning
  • On MATH and AIME benchmarks, it surpasses GPT-5 Turbo by significant margins
  • Available via API and in Gemini Advanced starting today

Google DeepMind today unveiled Gemini Reasoning 1.0, a specialized large language model designed for complex mathematical and scientific problem-solving. The company claims the model sets a new state-of-the-art on several leading benchmarks.

On the MATH benchmark — a standard test of mathematical reasoning across competition-level problems — Gemini Reasoning 1.0 scores 94.2%, compared to GPT-5 Turbo's 91.8%. On AIME 2026, the prestigious high school math competition, the model solved 28 of 30 problems, a result DeepMind described as "approaching the performance of the top 1% of human competitors."

The model uses an extended "thinking" mode that allocates more compute to difficult problems, similar to OpenAI's o-series reasoning models. Users can toggle between standard and extended reasoning modes depending on the complexity of their task.

Demis Hassabis, CEO of DeepMind, said the model represents "a significant step toward AI that can independently advance scientific discovery." The company plans to integrate the reasoning capabilities into future versions of Gemini for general use.

What This Means For You

If you work in STEM fields — engineering, data science, finance, or research — this is a model worth evaluating. Its ability to work through complex multi-step problems reliably makes it a strong candidate for technical workflows where GPT-5 sometimes makes arithmetic errors.

For general users, the improvements are more incremental. Reasoning-focused models tend to be slower and more expensive than general-purpose models for everyday tasks like writing or summarization. Use the right tool for the job.

Sources

Related News