understanding-ladder · demo

Why does latency explode once a server passes ~80% utilization?为什么服务器利用率过了 80%,延迟会突然暴涨?

One explanation, rewritten at every rung of Karpathy’s understanding ladder. Switch tabs to compare.同一段内容,按 Karpathy 的「理解阶梯」从第 0 级重写到第 4 级。点上面的标签切换。

Rung 0: what a model usually writes with no constraints. Mostly correct, but you read it twice.第 0 级:模型不加约束时的常见写法。内容大体正确,但要读两三遍。

Great question! It’s worth noting that, from a queueing-theory perspective, server latency exhibits a highly non-linear relationship with utilization. Specifically, when utilization is relatively low, incoming requests can typically be served promptly and queueing delay is largely negligible; however, as utilization climbs past a certain critical range (commonly cited as somewhere around 70–80%), bursty congestion caused by the inherent randomness of the arrival process becomes significantly amplified, causing average waiting time to grow sharply in a manner roughly proportional to 1/(1−ρ), which is precisely why operations teams in real-world production environments tend to provision a degree of headroom rather than maximizing resource utilization. Overall, understanding this is of great importance for capacity planning.

这是一个非常好的问题!值得注意的是,在排队论的视角下,服务器延迟与利用率之间呈现出一种高度非线性的关系。具体而言,当系统利用率处于较低水平时,请求到达后通常能够被及时处理,排队等待时间相对可以忽略不计;然而随着利用率逐步攀升并超过某个临界区间(通常认为在 70%–80% 左右),由于请求到达过程本身存在的随机性所导致的突发性拥塞会被显著放大,进而使得平均等待时间以一种近似于 1/(1−ρ) 的方式急剧增长,这也就是为什么在实际的生产环境中,运维团队往往会倾向于为系统预留一定的容量冗余,而不是追求将资源利用率最大化。总的来说,理解这一点对于容量规划具有重要的指导意义。

Rung 1: about 80% of the way to ASD-STE100. Nothing cut — the sentences are just clean.第 1 级:用 plain-chinese 重写。内容一点没删,只把句子拆干净。

Takeaway: The busier a server is, the longer requests wait. The growth is not steady: past about 80% utilization, waiting time rises sharply.

Two terms

  • Utilization (ρ): the fraction of time the server is busy. ρ = 0.8 means it serves requests 80% of the time.
  • Latency: the time from a request’s arrival until it is served. Latency = waiting time + service time.

Why it explodes

Requests do not arrive evenly. Sometimes many arrive in one second; sometimes none arrive. When requests arrive in a burst, the later ones wait.

If the server is often idle, the queue clears fast. If the server is rarely idle, the queue has not cleared when the next burst arrives. The queue grows.

In the simplest queueing model (M/M/1), average latency is:

latency = service time ÷ (1 − ρ)

The divisor (1 − ρ) is the fraction of time the server is idle. As ρ approaches 1, the divisor approaches 0, and latency goes toward infinity.

The numbers

Assume one request takes 10 ms to serve, on average:

Utilization ρAverage latency× idle latency
50%20 ms2×
80%50 ms5×
90%100 ms10×
95%200 ms20×

From 50% to 80%, utilization rises by 30 points and latency rises by 30 ms. From 90% to 95%, utilization rises by only 5 points, but latency rises by 100 ms.

So

Operations teams usually keep headroom. They do not push utilization to the maximum.

结论:服务器越忙,请求排队越久。而且不是匀速变久——利用率过了 80% 左右,排队时间会急剧变长。

两个术语

  • 利用率(ρ):服务器忙碌的时间占总时间的比例。ρ = 0.8 表示 80% 的时间在处理请求。
  • 延迟:一个请求从到达到处理完的总时间,等于「排队时间 + 处理时间」。

为什么会暴涨

请求不是均匀到达的。有时一秒来很多,有时一个也不来。请求扎堆来时,后到的请求要排队。

服务器空闲多时,队伍很快就能清空。服务器空闲少时,上一波队伍还没清空,下一波又来了。队伍越积越长。

在最简单的排队模型(M/M/1)里,平均延迟是:

延迟 = 处理时间 ÷ (1 − ρ)

分母 (1 − ρ) 是服务器的空闲比例。ρ 接近 1 时,分母接近 0,延迟就冲向无穷大。

算一下

假设处理一个请求平均要 10 毫秒:

利用率 ρ平均延迟比空闲时慢几倍
50%20 毫秒2 倍
80%50 毫秒5 倍
90%100 毫秒10 倍
95%200 毫秒20 倍

从 50% 到 80%,利用率多了 30 个百分点,延迟多了 30 毫秒。从 90% 到 95%,利用率只多了 5 个百分点,延迟却多了 100 毫秒。

所以

运维团队通常给服务器留出余量,不把利用率推到最高。

Rung 2: a diagram. Structure that text explains sentence by sentence is laid out at once: left, why requests queue; right, how long they wait.第 2 级:图。文字要一句一句讲的结构,图一次摊开:左边是「为什么排队」,右边是「排多久」。

① Random arrivals queue up in bursts① 请求随机到达,扎堆时排队 random arrivals随机到达 queue队列 server服务器 one at a time一次处理一个 latency = waiting + service延迟 = 排队时间 + 处理时间 often idle → queue clears fast空闲时间多 → 队伍很快清空 rarely idle → next burst arrives first空闲时间少 → 上一波没清空,下一波又来 → the queue keeps growing→ 队伍越积越长 ② latency = service ÷ (1 − ρ)② 延迟 = 处理时间 ÷ (1 − ρ) utilization ρ →利用率 ρ → latency延迟

How to read it: on the left, follow one request from left to right. On the right, the curve is average latency vs utilization (10 ms service time). It is nearly flat until 80%, then almost vertical.怎么读:左边从左往右,是一个请求经过的路径。右边的曲线是平均延迟随利用率的变化(处理时间 10 毫秒)。曲线在 80% 之前很平,之后几乎竖直向上。

Rung 3: a page. With text you trust the author; with a slider you see where latency takes off.第 3 级:网页。读文字时只能相信作者;拖动滑块时,你亲眼看到延迟在哪里开始暴涨。

Average latency平均延迟
of which waiting其中排队
× idle latency比空闲时慢

The animation is a random simulation: requests arrive at the current utilization; orange dots are the queue. Drag ρ above 0.9 and watch it grow.下面的动画是一个随机模拟:请求按当前利用率随机到达,橙色是队伍。把 ρ 拖到 0.9 以上,看队伍怎么变长。

Try this试试这样拖

  1. Drag slowly from 0.5 to 0.8. How much did latency grow?从 0.5 慢慢拖到 0.8,看延迟涨了多少。
  2. Go from 0.9 to 0.95: only 5 more points, and latency doubles.从 0.9 拖到 0.95,只多 5 个百分点,延迟翻倍。
  3. Change the service time: the curve keeps its shape and only scales up or down.改「平均处理时间」:曲线形状不变,只是整体变高或变低。

Rung 4: an explainer video. understanding-ladder does not render video; it delivers the storyboard, the toolchain, and a prompt to hand to a video tool.第 4 级:讲解视频。understanding-ladder 不渲染视频,只交付文字稿、分镜、工具链和一段可以直接交给视频工具的 prompt。

Storyboard (~75 s)分镜(约 75 秒)

#LengthVisualNarration
15 sTitle cardThis video explains one thing: why a server slows down sharply past 80%.
210 sBlue dots drop into a queue; the server takes one at a timeRequests arrive at random. The server handles one at a time.
312 sρ = 0.5: a queue forms and clearsThe server is idle half the time. A queue forms now and then, and clears fast.
412 sρ = 0.9: the queue keeps growingThe server is busy 90% of the time. The next burst arrives before the queue clears.
514 sCurve 1/(1−ρ) drawn on the right; a dot slides to 0.95Average latency is service time divided by idle time. As idle time nears zero, latency goes to infinity.
612 sFour rows appear: 2×, 5×, 10×, 20×With 10 ms service time, latency goes from 20 ms at 50% to 200 ms at 95%.
710 sRecap cardTo recap: latency depends on idle time. Keep headroom.
#时长画面旁白
15 秒标题卡这段视频讲一件事:服务器为什么在 80% 之后突然变慢。
210 秒蓝点随机落进一条队列,服务器一次处理一个请求随机到达。服务器一次只能处理一个。
312 秒ρ = 0.5:队伍时有时无服务器一半时间空闲。队伍偶尔出现,很快清空。
412 秒ρ = 0.9:队伍不断变长服务器九成时间在忙。上一波还没清空,下一波又来了。
514 秒右侧画出曲线 1/(1−ρ),一个点沿曲线滑到 0.95平均延迟等于处理时间除以空闲比例。空闲比例接近零,延迟冲向无穷大。
612 秒四行数字依次出现:2 倍、5 倍、10 倍、20 倍处理时间 10 毫秒时,从 50% 到 95%,延迟从 20 毫秒涨到 200 毫秒。
710 秒回顾卡回顾:延迟取决于空闲比例。给服务器留余量。

Toolchain工具链

  • showtime: open source; your coding agent renders video locally. Since 0.4.0 it reads this storyboard table directly (--from-storyboard).开源,让 coding agent 在本机渲染视频。0.4.0 起能直接读取这张分镜表(--from-storyboard)。
  • Manim Community: the 3Blue1Brown-style math animation library.3Blue1Brown 风格的数学动画库。
  • kokoro-onnx: free local voice. Or ElevenLabs (API key needed).免费本地配音。或用 ElevenLabs(需要 API key)。

Prompt for the video tool交给视频工具的 prompt

Create a 3b1b style video explainer, about 75 seconds, on why server latency explodes past ~80% utilization. Follow this storyboard exactly: […]. Free local TTS voice. Burn in captions.

Easier to understand ≠ verified. The ladder lowers the cost of understanding, not the cost of checking. This page uses the simplest M/M/1 model: exponential arrivals and service times, one server. Real systems give different numbers, but “latency rises sharply near full load” holds broadly. 好懂 ≠ 已验证。阶梯降低的是理解成本,不是验证成本。本页用的是最简单的 M/M/1 模型:到达和处理时间都服从指数分布,只有一台服务器。真实系统的数字会不同,但「接近满载时延迟急剧上升」这个规律普遍成立。