MIT Researchers Hit 96 Percent on MATH Benchmark with New Reasoning Model
MIT researchers published a new reasoning architecture on July 15 achieving 96 percent on the MATH benchmark. The model uses symbolic verification layers combined with transformer backbones. Results surpass previous state-of-the-art by 11 percentage points.
Stanford Researchers Achieve 94 Percent on GPQA with New Architecture
Stanford University researchers published a new architecture on July 11 that scores 94 percent on GPQA. The model uses dynamic sparse attention and requires 60 percent less compute than dense transformers. The paper details training on 2.1 trillion tokens.
Google DeepMind Publishes AlphaCode 3 Achieving 87 Percent on Codeforces
Google DeepMind released AlphaCode 3 research results on July 10 2026 scoring 87 percent on Codeforces contests. The system solves problems at the level of top 5 percent human competitors. It matters because it marks the first time an AI system consistently outperforms elite human programmers in competitive settings.