Meta released Llama 4 on July 22 2026 with a 10 million token context window and native video processing capabilities. The model achieves 92.4 percent on MMLU and 89.1 percent on VideoMME. Training used 2.8 trillion tokens across 128,000 H100 GPUs completed on July 18.
Developers can access Llama 4 through the new Llama API starting July 25 with rate limits of 1,000 requests per minute for free tier users. Meta also open sourced the 70B parameter variant under the same license as Llama 3.1.
Background work began in January 2025 when Meta expanded its AI research team to 4,200 employees. The company invested 18 billion dollars in compute infrastructure during fiscal 2025.
Previous Llama releases established Meta as the leader in open weights models with over 500 million downloads of Llama 3.1 by June 2026.
Why this matters
Llama 4 gives enterprises a high-performance open alternative to closed models from OpenAI and Anthropic. Its 10 million token context enables new applications in legal document analysis and long-form video summarization that were previously cost prohibitive.
Competitors including Google and Mistral are expected to respond with larger context windows within 90 days according to industry analysts.
Meta plans to release Llama 4.1 with improved agentic capabilities by October 2026 and will expand API availability to additional regions in August.