The complaint
"It was fine for 45 minutes, then it started choking."
That is the whole bug report, and it is a good one. It has a time, a before, and an after. It also has a built-in theory that almost everyone reaches for, myself included: something must expire at 45 minutes.
That theory is wrong, and the way it is wrong is worth an article.
Why 45 minutes feels like a timer
When a system runs clean and then degrades at a repeatable-sounding moment, the mind goes straight to expiry. Sessions expire. Signed links expire. Tokens expire. Free tiers cut off. Every one of those produces a genuine, sharp, at-the-clock failure, so the instinct is trained by real experience.
So I checked. Every timer in the path got written down and compared against the failure moment: the signing window on delivery links, the login session, the storage retention policy, the reconnect intervals. None of them landed at 45 minutes. Not close. The nearest one was an hour away, and it fails differently... a hard stop, not a slow choke.
That is the first useful lesson. A timer failure is a cliff. This was a slope. The stream did not stop; it got progressively worse. Slopes and cliffs have different causes, and the shape of the failure tells you which one you are hunting before you know anything else.
Measuring instead of guessing
The next move was the one that actually solved it: stop reasoning and go measure.
A live stream is not one object. It is a long series of small media segments, produced continuously, uploaded continuously, and fetched continuously by whoever is watching. Every one of those segments carries a timestamp for when it arrived at the delivery layer. I had every segment from the session still on disk.
So instead of arguing with myself, I swept them. For each segment, compare when it should have arrived against when it actually did, and plot the difference minute by minute. That produces a lag curve... a single line that shows the health of the entire session in one shape.
The curve told the story immediately:
- For the first stretch, flat. Segments arriving essentially on schedule.
- Then, spikes. A few seconds late, recovering, late again. Not fatal... a viewer might not notice, but the margin was being eaten.
- Then a long stall. Nothing arrived for a substantial gap.
- After that, permanently degraded. Never flat again.
The choke the operator felt at "45 minutes" was the end of that sequence, not the start. The failure began much earlier, in the small spikes nobody could see, because at that stage the system was still absorbing them.
The part that surprised me
Here is the mechanism, and it is the reason this article exists.
Live streaming keeps a rolling window of recent segments available and discards older ones. It has to... otherwise storage grows without bound during a long service. That window is generous relative to normal operation, which is exactly why it is invisible when things are healthy.
Now put a stall inside that window. If a segment is late by less than the window, it arrives to find its place still open. Everything continues. If a segment is late by more than the window, it arrives to find its slot already cleaned up. That segment is gone. Not delayed... gone. The viewer's player hits a hole, and the recovery from a hole is far more expensive than the recovery from a delay.
So the sequence is:
- Upload jitter causes small delays. System absorbs them.
- Jitter accumulates. Delays grow.
- One delay exceeds the rolling window. A segment is lost.
- The player, now recovering from a gap, is less able to absorb the next delay.
- Repeat.
That is not a timer. That is a feedback loop, and feedback loops always look like they start suddenly, because they do... from the outside. From the inside they were building the entire time.
What was actually wrong
The upload path from the venue was jittery. Not slow... jittery. Those are different problems, and confusing them costs people a lot of money.
A slow connection is easy to diagnose. You run a speed test, see a low number, and buy more bandwidth. A jittery connection tests fine. The average is healthy. The peak is healthy. But the delivery is uneven: bursts, pauses, bursts. Averaged over a minute it looks great. Measured segment by segment, it is a sawtooth.
Live streaming does not care about your average. It cares whether this segment gets out before the next one is ready. A connection that delivers a minute's worth of data in fifty-five seconds of transfer and five seconds of silence has excellent throughput and can still break a stream.
Worth saying plainly: my first theory was that the cloud side was at fault. The segment timings cleared it. I have been wrong about a cause often enough to measure before I act on a suspicion, and this is one of the times that habit paid for itself.
What I changed
Two things, and the order matters.
First, I widened the tolerances. The rolling window and the retry behaviour were both tuned for a clean network, which is a reasonable default and a bad assumption. Widening them costs a little memory and buys a lot of forgiveness. A stall that used to lose a segment now gets absorbed.
Second, and this is the one I consider the real fix, I made the failure legible. The old behaviour degraded silently. The operator's only signal was a viewer complaint, which arrives late and vague. Now the condition surfaces where the operator can see it, while there is still time to act.
That second change fixes no bug. It changes who is in a position to notice.
The lesson I keep relearning
I have written some version of this lesson down more than once, so here it is plainly.
When something fails after running fine, the interesting question is not "what happened at that moment"... it is "what was accumulating before it." The visible failure is where the accumulation crossed a threshold. Investigating the moment finds the threshold. Investigating the slope finds the cause.
The corollary, for anyone running a live day: if your stream degrades over a long service rather than failing outright, look at your upload path's consistency before you look at its speed. And be suspicious of any explanation that requires a coincidence... a stream that fails at a repeatable clock time has a timer; a stream that fails at a vague "about 45 minutes" has a slope.
I got a better stream out of this. I also got a better instinct, which lasts longer.

