LaunchDevelopersJuly 8, 2026

Open, convenient and predictable: Introducing Provisioned Throughput

Provisioned Throughput provides reserved inference capacity for open models like MiniMax M3 and GLM-5.2 with a 99% uptime SLA and token-based pricing, costing up to 90% less than Claude Opus 4.8. Together AI reports API token volume grew from 30 billion to 400 trillion tokens per month over nine months.

1 source