It was fine for 45 minutes

A stream that ran clean then started choking. I blamed the cloud, measured it, and found the fault was 45 minutes upstream of where it appeared.

By Abelitie · September 25, 2026
It was fine for 45 minutes

The complaint

"It was fine for 45 minutes, then it started choking."

That is the whole bug report, and it is a good one. It has a time, a before, and an after. It also has a built-in theory that almost everyone reaches for, myself included: something must expire at 45 minutes.

That theory is wrong, and the way it is wrong is worth an article.

Why 45 minutes feels like a timer

When a system runs clean and then degrades at a repeatable-sounding moment, the mind goes straight to expiry. Sessions expire. Signed links expire. Tokens expire. Free tiers cut off. Every one of those produces a genuine, sharp, at-the-clock failure, so the instinct is trained by real experience.

So I checked. Every timer in the path got written down and compared against the failure moment: the signing window on delivery links, the login session, the storage retention policy, the reconnect intervals. None of them landed at 45 minutes. Not close. The nearest one was an hour away, and it fails differently... a hard stop, not a slow choke.

That is the first useful lesson. A timer failure is a cliff. This was a slope. The stream did not stop; it got progressively worse. Slopes and cliffs have different causes, and the shape of the failure tells you which one you are hunting before you know anything else.

Measuring instead of guessing

The next move was the one that actually solved it: stop reasoning and go measure.

A live stream is not one object. It is a long series of small media segments, produced continuously, uploaded continuously, and fetched continuously by whoever is watching. Every one of those segments carries a timestamp for when it arrived at the delivery layer. I had every segment from the session still on disk.

So instead of arguing with myself, I swept them. For each segment, compare when it should have arrived against when it actually did, and plot the difference minute by minute. That produces a lag curve... a single line that shows the health of the entire session in one shape.

The curve told the story immediately:

  • For the first stretch, flat. Segments arriving essentially on schedule.
  • Then, spikes. A few seconds late, recovering, late again. Not fatal... a viewer might not notice, but the margin was being eaten.
  • Then a long stall. Nothing arrived for a substantial gap.
  • After that, permanently degraded. Never flat again.

The choke the operator felt at "45 minutes" was the end of that sequence, not the start. The failure began much earlier, in the small spikes nobody could see, because at that stage the system was still absorbing them.

The part that surprised me

Here is the mechanism, and it is the reason this article exists.

Live streaming keeps a rolling window of recent segments available and discards older ones. It has to... otherwise storage grows without bound during a long service. That window is generous relative to normal operation, which is exactly why it is invisible when things are healthy.

Now put a stall inside that window. If a segment is late by less than the window, it arrives to find its place still open. Everything continues. If a segment is late by more than the window, it arrives to find its slot already cleaned up. That segment is gone. Not delayed... gone. The viewer's player hits a hole, and the recovery from a hole is far more expensive than the recovery from a delay.

So the sequence is:

  1. Upload jitter causes small delays. System absorbs them.
  2. Jitter accumulates. Delays grow.
  3. One delay exceeds the rolling window. A segment is lost.
  4. The player, now recovering from a gap, is less able to absorb the next delay.
  5. Repeat.

That is not a timer. That is a feedback loop, and feedback loops always look like they start suddenly, because they do... from the outside. From the inside they were building the entire time.

What was actually wrong

The upload path from the venue was jittery. Not slow... jittery. Those are different problems, and confusing them costs people a lot of money.

A slow connection is easy to diagnose. You run a speed test, see a low number, and buy more bandwidth. A jittery connection tests fine. The average is healthy. The peak is healthy. But the delivery is uneven: bursts, pauses, bursts. Averaged over a minute it looks great. Measured segment by segment, it is a sawtooth.

Live streaming does not care about your average. It cares whether this segment gets out before the next one is ready. A connection that delivers a minute's worth of data in fifty-five seconds of transfer and five seconds of silence has excellent throughput and can still break a stream.

Worth saying plainly: my first theory was that the cloud side was at fault. The segment timings cleared it. I have been wrong about a cause often enough to measure before I act on a suspicion, and this is one of the times that habit paid for itself.

What I changed

Two things, and the order matters.

First, I widened the tolerances. The rolling window and the retry behaviour were both tuned for a clean network, which is a reasonable default and a bad assumption. Widening them costs a little memory and buys a lot of forgiveness. A stall that used to lose a segment now gets absorbed.

Second, and this is the one I consider the real fix, I made the failure legible. The old behaviour degraded silently. The operator's only signal was a viewer complaint, which arrives late and vague. Now the condition surfaces where the operator can see it, while there is still time to act.

That second change fixes no bug. It changes who is in a position to notice.

The lesson I keep relearning

I have written some version of this lesson down more than once, so here it is plainly.

When something fails after running fine, the interesting question is not "what happened at that moment"... it is "what was accumulating before it." The visible failure is where the accumulation crossed a threshold. Investigating the moment finds the threshold. Investigating the slope finds the cause.

The corollary, for anyone running a live day: if your stream degrades over a long service rather than failing outright, look at your upload path's consistency before you look at its speed. And be suspicious of any explanation that requires a coincidence... a stream that fails at a repeatable clock time has a timer; a stream that fails at a vague "about 45 minutes" has a slope.

I got a better stream out of this. I also got a better instinct, which lasts longer.

Asked often.

Why did my stream start buffering partway through instead of at the start?

Buffering that appears late is usually accumulated upload jitter, not a timer... small hiccups build a backlog until the delivery window can no longer absorb them.

Does a stream have a time limit that causes it to fail?

Not normally. If a stream fails at a consistent clock time suspect an expiring credential; if it degrades gradually, suspect your upload path.

My stream went bad mid-service and I had no idea until someone told me. Is there anything I can watch?

Yes, the console now surfaces the condition while it is still building rather than after viewers notice. If you see it drifting, dropping to a lower quality mid-service is the cheapest recovery.

Our internet speed test looks fine. Does that mean streaming will be fine?

Not necessarily, and this is the trap. A connection can average well and still deliver unevenly, and live streaming cares whether this second's data got out rather than what the minute averaged.

You blamed your cloud provider first. How much of this diagnosis should I trust?

Trust the measurement rather than the diagnosis. I was wrong about the cause until I swept every segment's arrival time, and that curve is the evidence the conclusion rests on.

Why were the tolerances set for a clean network to begin with?

Because a clean network is what I had. It is a reasonable default and a bad assumption, and widening it costs a little memory in exchange for a lot of forgiveness.

What does the software do on its own when the upload gets choppy?

It absorbs a wider range of stalls than it used to before anything is lost, keeps retrying rather than giving up, and shows the operator the condition instead of degrading quietly.

Try it on your Sunday.

Free tier, no credit card. A laptop and the phones in the room.

Start Free →
It was fine for 45 minutes | Broadcasteer