🔬 Research / /via The Verge / updated 6d ago

Stanford Researchers Achieve 94 Percent on GPQA with New Architecture

Stanford University researchers published a new architecture on July 11 that scores 94 percent on GPQA. The model uses dynamic sparse attention and requires 60 percent less compute than dense transformers. The paper details training on 2.1 trillion tokens.

#StanfordUniversity
~/ Research/ Stanford Researchers Achieve 94 Percent on GPQA...

Stanford researchers released a new transformer variant on July 11 achieving 94 percent on the GPQA benchmark. The architecture employs dynamic sparse attention that activates only 12 percent of parameters per token. Training used 2.1 trillion tokens over 18 days on 9,000 H100 GPUs.

Energy consumption dropped to 1.8 megawatt-hours per training run compared with 4.5 for equivalent dense models. The team open-sourced weights and training code under an Apache 2.0 license. Inference speed improved 3.2x on A100 hardware.

GPQA tests graduate-level science questions across physics, chemistry, and biology. Prior state-of-the-art models reached 89 percent in May 2026. The Stanford approach builds on mixture-of-experts research from 2024.

Why this matters

The result narrows the gap between academic and frontier lab capabilities. Researchers gain efficient alternatives to closed models for scientific discovery. This may accelerate open-source progress in specialized domains.

Industry labs will likely adopt sparse attention patterns in next-generation systems. Energy efficiency gains support sustainability goals for large-scale training. The open release enables rapid community validation and extension.

Follow-up work will target multimodal GPQA variants by September. Stanford plans a public leaderboard update within two weeks.

share
𝕏 FB
← cd ../news