🛠️ Tools / /via VentureBeat / updated -107m ago

Hugging Face Launches Inference Endpoints 2.0 with Serverless GPUs

Hugging Face introduced Inference Endpoints 2.0 on July 26 2026 offering serverless GPU inference. The platform auto-scales from zero to thousands of replicas with sub-second cold starts. It matters because it lowers the barrier for deploying production-grade models.

#HuggingFace
~/ Tools/ Hugging Face Launches Inference Endpoints 2.0 w...

Hugging Face launched Inference Endpoints 2.0 on July 26 2026 providing serverless GPU inference for open models. The service supports NVIDIA H100 and AMD MI300X accelerators with automatic scaling from zero replicas. Cold-start latency averages 800 milliseconds for typical transformer workloads.

Pricing starts at 0.0004 dollars per second on H100 hardware with volume discounts for sustained usage. The update includes built-in metrics, canary deployments, and integration with the Hugging Face Hub for one-click publishing.

The company previously offered dedicated endpoints requiring manual instance management. Serverless removes the need for capacity planning while maintaining the same security isolation guarantees.

Why this matters

The release accelerates the shift toward consumption-based model deployment and reduces operational overhead for startups and research teams. It also strengthens Hugging Face's position as the default platform for open-source model distribution.

Competitors including Replicate and Together AI are expected to respond with similar offerings. Hugging Face plans to add TPU support and multi-region failover before the end of 2026.

share
𝕏 FB
← cd ../news