AnalysisDevelopersSeptember 19, 2026

OpenAI engineers detail LLM inference routing in production

OpenAI's inference load balancer originally set routing weights via a proportional controller: engines reported signals, a controller smoothed them into a score, compared it to the fleet average, and nudged each weight up or down.

People · Qianru Lao, Lu Zhang

1 source

More stories today

Open the live feed