OpenAI's Codex Spark achieves 1,000 tokens per second on Cerebras

The GPT 5.3 Codex Spark model reached 1,000 tokens per second on Cerebras hardware, shifting the primary inference bottleneck from compute to network latency. To address this, the team implemented a persistent websocket mode to maintain stateful context and reduce overhead compared to standard HTTP server-sent events.
Featured · Dominik Kundel
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Paper examines the limitations of current AI evaluation methods
- OnlyHuman filter list removes AI-generated SEO spam from search results
- Qwen tokenizes 330-line code into 1,609 tokens; Gemma needs 4,258
- LifeOS: open-source AI harness for personal growth and work
- MINIMAX video drops Indiana Jones into Mortal Kombat