OpenAI reduces GPT-5.6 inference costs by 20% via self-optimization

OpenAI reports a 20% reduction in serving costs for GPT-5.6 by using the model to autonomously rewrite production kernels in Triton and Gluon. Additionally, speculative decoding improvements have increased token-generation efficiency by over 15%.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Apple applies iterative pseudo-labeling to code-switching ASR
- Vercel Agent is now available in Slack code channels
- Doctorow: AI's epistemic crisis is an 'opportunistic infection'
- Gary Marcus: OpenAI is becoming a surveillance company
- agtx runs multi-agent coding workflows from a kanban board