TL;DR

Google has released Gemini 3.1 Pro, which scored 77.1% on the ARC-AGI-2 benchmark — more than double the reasoning performance of Gemini 3 Pro. The model is rolling out across the Gemini app, Google AI Studio, Vertex AI, NotebookLM, and the Gemini CLI.

What Happened

Gemini 3.1 Pro builds on last week’s update to Gemini 3 Deep Think, which targeted science, research, and engineering tasks. Google described 3.1 Pro as the upgraded core intelligence behind those capabilities, now being shipped across consumer and developer products.

The model is available in preview for developers via the Gemini API in Google AI Studio, Gemini CLI, Google Antigravity (the company’s agentic development platform), and Android Studio. Enterprise users can access it through Vertex AI and Gemini Enterprise. Consumers get it via the Gemini app and NotebookLM, with higher limits for AI Pro and Ultra plan subscribers.

Google highlighted several practical applications: generating website-ready animated SVGs from text prompts, building a live aerospace dashboard by configuring public telemetry streams, coding interactive 3D visualisations with hand-tracking support, and translating literary themes into functional web designs.

Why It Matters

The ARC-AGI-2 benchmark measures a model’s ability to solve entirely new logic patterns — tasks it hasn’t seen before. Doubling performance on this benchmark suggests a substantial step forward in general reasoning capability rather than improvements on familiar task types.

The broad rollout across Google’s ecosystem — from consumer apps to enterprise platforms and developer tools — means the upgraded reasoning is immediately available for production use rather than remaining a research preview.

Looking Forward

Google said 3.1 Pro is in preview while the company validates updates and advances agentic workflow capabilities. A general availability release is expected soon. The rapid iteration — Gemini 3 Pro launched in November, with 3.1 Pro arriving just three months later — reflects an accelerating release cadence in the model development race.