AnalysisDevelopersSeptember 19, 2026

CoreWeave's Sitanshu Gupta on scaling inference from MVP to trillion-parameter workloads

On agentic requests, 80-90% of input tokens are identical to the prior request, making prefill the most expensive operation in an inference stack — which is why cached input tokens are priced far below fresh ones.

People · Sitanshu Gupta

1 source

More stories today

Open the live feed