Amazon Web Services introduced Trainium3 chips on July 16 2026 with four times the training throughput of the prior generation while maintaining similar power consumption. The chips feature 128 GB of high-bandwidth memory and a new collective communication fabric. AWS will make Trainium3 available in UltraCluster instances beginning August 4 2026.
Internal benchmarks show a 2.8 trillion parameter model trains 35 percent cheaper per token compared with equivalent GPU clusters. AWS also released updated Neuron SDK 3.0 with automatic parallelism for the new silicon.
The launch follows AWS's 2025 investment of 10 billion dollars in custom silicon fabrication partnerships. Trainium3 targets both internal Amazon teams and external customers seeking GPU alternatives.
Early access customers include several large media and pharmaceutical companies already running production workloads on Trainium2.
Why this matters
Trainium3 strengthens AWS's position in the custom AI accelerator market and reduces reliance on NVIDIA supply constraints. Lower training costs could accelerate model iteration cycles for mid-sized AI labs.
Industry observers expect Google and Microsoft to respond with their own next-generation TPU and Maia announcements before the end of 2026.
Broader availability may shift more training workloads onto cloud-specific silicon over the next 18 months.