OpenAI launched the o4 reasoning model on July 22 2026 with a 4 million token context window and native support for multi-step tool use. Internal benchmarks show 92.4 percent on GPQA Diamond and 89.1 percent on SWE-Bench Verified. The release includes a new API tier priced at 15 dollars per million input tokens.
Developers can access o4 through the existing Chat Completions endpoint with a dedicated reasoning parameter. Early testers at Scale AI and Adept reported 40 percent fewer hallucinated steps on complex coding tasks. OpenAI stated the model was trained on a mixture of synthetic reasoning traces and verified code repositories through June 2026.
Background includes OpenAI's prior o-series releases that emphasized test-time compute scaling. The company has maintained a monthly cadence of capability updates since early 2025 while competitors focused on larger context alone.
Why this matters
The o4 launch intensifies pressure on Anthropic and Google to demonstrate comparable reasoning gains rather than context length alone. Enterprises evaluating AI for software engineering workflows now have clearer performance data on multi-step tasks.
Analysts expect o4 to accelerate adoption of agentic systems in legal and financial services where step-by-step verification is required. OpenAI plans to release a distilled o4-mini variant by September 2026.