LaunchDevelopersJuly 8, 2026

Together AI launches Provisioned Throughput for open models

The new inference service offers reserved capacity for MiniMax M3 and GLM-5.2 with a 99% uptime SLA and token-based pricing. It claims to reduce inference costs by up to 90% compared to Claude Opus 4.8.

4 sources

MiniMax by email

Get an email when MiniMax ships something

More stories today

Open the live feed