OpenAI's Jalapeño chip beats Nvidia Blackwell on inference efficiency

OpenAI published first benchmarks for Jalapeño, its Broadcom-built inference chip: 1.5–1.9× more AI work per watt and 1.7–3.6× lower end-to-end latency than Nvidia GB200/GB300 systems. Tested on SemiAnalysis's InferenceX benchmark across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.
People · Richard Ho
How this story unfolded
4 days · 10 reports · 11 community posts · 21 of 24 shown
- Aug 25
OpenAI says its Jalapeño chip can power faster AI responses than the competitiontheverge.com
OpenAI built a chip in nine months. Then it let AI rewrite the code.thenewstack.io
OpenAI Claims Its New Chips Can Outperform Nvidia Processors in Testsbloomberg.com
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks showtechcrunch.com
Jalapeño’s first results show industry-leading speed and efficiency in AI inferenceopenai.com
OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testingbloomberg.com
- Aug 26
- Aug 27
- Aug 28
- Aug 29
More stories today
datasette-explain 0.2.2 adds explain plans to read-only stored-query pages
datasette-explain 0.2.2 ships explain plans on read-only stored-query pages, released after Simon Willison upgraded datasette.simonwillison.net to Datasette 1.0a40.
Simon Willison's Weblog·4 hours ago
Emily Bender critiques AI's "freeing up time" selling point
Emily M. Bender·4 hours agofocus-llama fork brings Declarative Attention to llama.cpp
A llama.cpp fork implements Declarative Attention from arXiv:2609.02737 (Google DeepMind and KAIST AI), letting the model declare in its own output which context chunks it needs while the engine restricts what following tokens attend to. No scorer and no training are required — just prompting.
r/LocalLLaMA·5 hours ago
Google's AX orchestrator runs agentic tasks at cluster scale
AX declares agentic tasks in YAML (Workspace and Task kinds, apiVersion ax.io/v1alpha1) and runs them in sandboxes with CPU/memory limits, network fencing, and Git/MCP workspace setup. Tasks can be suspended, resumed, and deleted via the ax CLI.
Hacker News·5 hours agoDemis Hassabis says creativity separates good from great scientists
In a Royal Society of Arts conversation with Professor Hannah Fry, the Google DeepMind co-founder argues creativity and intuition are fundamental to scientific discovery and breakthrough ideas.
Royal Society of Arts·5 hours ago
Reddit user builds DIY Jev-like inference setup with open models
A r/LocalLLaMA user describes experimenting with a simple Jev-like inference setup using ordinary open models, calling the Jev trend overhyped by a large margin.
r/LocalLLaMA·5 hours ago
Reddit debates $1k Taalas chip running Qwen3.8-27B at 7,000 TPS
A r/LocalLLaMA thread asks whether users would buy a hypothetical Taalas consumer chip for $1,000 that runs Qwen3.8-27B at 7,000 tokens per second. The poster notes the chip would be locked to that model but would not become obsolete when Qwen4-27B launches.
r/LocalLLaMA·5 hours agoCohere AI platform automates task completion
Cohere·6 hours ago