Most AI work can wait: routing cuts costs
Tom Tunguz argues teams should design routing before choosing models, claiming 70-80% of traffic can run on local or async models, cutting AI spend by 90%+. Cites Coinbase halving AI spend while token usage grew.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Stability AI raises $232M backed by music and gaming giants
- OpenAI's Jalapeño chip beats Nvidia in inference benchmarks
- a16z podcast explores AI's impact on computing's evolution
- AI adoption lags in legal due to fragmented data foundations
- Bain & Company joins Claude Partner Network as Global Premier partner