Hacker News

realberkeaslan
Ask HN: What makes it so hard to keep LLMs online?

It feels like every few days one of the big AI services is down, degraded, or just slow. I don't mean this as a complaint. I'm just genuinely curious. These are well-funded companies with smart people. What is it about running these models that makes reliability so elusive? Is it just demand nobody predicted, or is there something fundamentally different about serving AI vs. a normal web app?


angarrido5 days ago

must people think it’s just GPU cost. In practice it’s coordination: model latency variance + queueing + retries under load. You don’t scale linearly, you get cascading slowdowns.

zipy1245 days ago

Likely one large contributor is that for a normal service, if it's down it's as simple as re-routing to another service, and there is basically an unlimited amount of CPU servers around the world to spin up on demand. GPU servers are much harder to spin up on demand, as supply is so constrained.

Another factor is just it's a new field and move fast and break things is still the go to as competition is high, and the stakes are incredibly high monetary wise.

A pessimistic, but perhaps true theory is also just vibe-coding/slop is reducing their reliability.

A counter point is that regular services like github seem to go down almost as frequently.

blemis5 days ago

honestly it's mostly gpu supply. scaling up to handle load means spinning up new nodes, and that takes minutes not seconds because the models are huge and need multiple coordinated gpus per instance.

also worth saying, even when things are "up" you often get different answers to the same question. that's the reliability problem nobody talks about. fine for a chatbot, not fine if you're building anything that needs to be repeatable and deterministic... i moved more to the ML route, but i guess it depends on what you are trying to do.

arc_light5 days ago

[dead]

andyjohnson05 days ago

Not sure why this was apparently flagged to death. Vouched.

roywiggins5 days ago

It's the em-dashes from a green account.

andyjohnson04 days ago

Account's comment history didn't look particularly AI generated. And, as an organic human who uses em dashes myself, I kind of hope we can get past this simplistic take that they are a signifier of ai content.

Besides that, I thought the comment had something useful to say — whether ai-generated or not.

moomoo114 days ago

Ai slop

If you can’t tell then damn idk man.

hn-front (c) 2024 voximity
source