Thanks for @‘ing me Tengyu - I don’t know for sure about OpenAI, but a few common challenges that happen to many inference scenarios:
- small unnoticed bottlenecks. There was one year of Alibaba’s double eleven event (equivalent of Black Friday, but much bigger) when people could not check out. Turned out the address normalization service, a tiny service in the whole chain, was overloaded. This brought down the end to end service. Other similar small services might be authentication, offensive word filtering, etc. I think this might be a possible reason for chatgpt going down.
- Traffic simply got overloaded. This isn’t common for microservices, as scaling only takes sub-seconds. LLMs are more prone to that, because loading models and doing other preprocessing takes minutes. The likelihood of this being the reason for chatgpt downtime is low, as I believe traffic won’t be so bursty with OpenAI’s volume.
- System components, like Kubernetes or load balancers, going wrong during regular updates or maintenance. This happens more often than people expect. A few months ago there was a company who updated Kunernetes a few major versions up without checking, and it wasn’t the best day for the CTO.
- a major part of service being disrupted due to things like power outage, leading to the other services flooded and overloaded. This also affects recovery as well, because poor servers who recover without its peers will face the whole angry awaiting traffic (think Jon Snow in the battle of bastards) and immediately go overload. Mature traffic control has to be in front of the services to reject any volume unsupportable by the current capacity. We at
@LeptonAI had a client who literally suffered from this: they accidentally shut down the main inference service manually during a Saturday. Our gateway saved the day and things were back in as short as 10 mins.
- unexpected global network outage. Yes this happens. There were a couple times when some network providers between us-east and us-west were down for a couple hours or even days this year. This is more often in neocloud providers, as hyper scalers normally have their own backbone network. We solved the problem by building a logical virtual network across all our multi-cloud servers, and we route traffic through normal providers / aws / gcp backup to ensure uptime.
My cofounder who was a former CNCF committee member took care of all these so I can be a happy clueless CEO. In summary: a series of seemingly boring but essential work behind the awesome models that researchers build. Combining these and you take off to stratosphere or even higher.