Google DeepMind released Gemini 3 on July 27 2026 with a native 10 million token context window and integrated video audio and code modalities. The model achieved 94.2 percent on MMLU 89.7 percent on GPQA and 82.4 percent on SWE-Bench Verified. Training used 12 trillion tokens across 16,000 TPU v5 pods completed on July 15.
Internal benchmarks show Gemini 3 reduces hallucination rates by 41 percent compared with Gemini 2.5 on long-document QA tasks. The model supports real-time tool calling across 47 APIs and runs inference at 240 tokens per second on Cloud TPU v6.
Google also opened a limited API preview for 500 enterprise customers starting August 1. Pricing starts at 0.003 dollars per thousand input tokens for contexts under 1 million tokens.
Google DeepMind has iterated on Gemini every nine months since the 2023 debut. The new architecture incorporates mixture-of-experts routing with 1.8 trillion active parameters out of 8.4 trillion total.
Competitors including OpenAI and Anthropic are expected to respond with updated context lengths before year end.
Why this matters
Gemini 3 sets a new standard for context length that enables entire codebases and multi-hour video analysis in a single prompt. Enterprises can now replace retrieval-augmented generation pipelines with direct long-context inference reducing latency and cost.
The performance jump accelerates adoption of AI agents for software engineering and scientific literature review. Regulators will likely scrutinize the model's expanded capabilities in upcoming safety evaluations.
Google plans to integrate Gemini 3 into Search and Workspace products by September 2026 expanding its reach to billions of users.