Why your lip sync drifts (and why it is almost never the camera)

Audio and video arrive at different times for structural reasons. Here is where the delay actually comes from, why it drifts mid-service, and how to fix it properly.

By Abelitie · September 29, 2026
Why your lip sync drifts (and why it is almost never the camera)

The problem everyone has and nobody enjoys

The mouth moves, then the word arrives. Or the word arrives first, which is worse... it reads as wrong immediately, in a way that a slight video delay does not.

Lip sync is one of the most common complaints in live production and one of the most frustrating to chase, because the usual advice is to nudge a slider until it looks right, which works until the next service, when it does not.

This is an article about where the delay actually comes from. Once you know that, the fix stops being guesswork.

Audio is fast, video is slow

Here is the fundamental asymmetry, and almost everything else follows from it.

Audio has a short path. A microphone produces a signal, the signal is digitised, and it is ready. There is processing, but it is comparatively cheap and comparatively quick.

Video has a long path. A camera sensor captures an image. That image is converted into a usable format. It may be scaled, colour-converted, composited with other layers... lyrics, logos, another camera. Then it is compressed, which is by far the most expensive step, because compression works by comparing frames to each other and therefore has to hold several frames before it can emit any.

Every one of those stages costs time. Add them up and video is meaningfully behind audio by the time both reach the output.

So the natural state of any live production system is audio-first. Not because anything is broken... because that is the shape of the work.

Therefore the correct fix is to delay the audio. You cannot make video faster; the stages are doing necessary work. You can hold the audio back until they match. Anyone who tells you to advance the video is describing something that is not physically available.

Where the delay is added matters enormously

This is the part that separates a fix that holds from a fix that unravels.

Alignment can be applied at two very different points, and they behave completely differently.

Downstream alignment is a correction applied near the end... at the player, or the platform, or as a display-side adjustment. It fixes what you can see right now. It does not change what was recorded, and it does not travel with the stream. Every destination needs its own correction, and your local recording, which never passed through that correction, is still wrong.

Upstream alignment happens before the audio and video are combined into a single stream. The alignment is baked in. Every downstream destination (the live stream, the recording, the archive, anything that gets the file later) inherits correct timing, because the timing is a property of the content rather than a setting on a device.

The practical implication: if you have fixed lip sync in more than one place, you have not fixed lip sync. You have applied several independent corrections that will drift apart from each other. Find the earliest point where audio and video meet, and fix it there once.

Why platforms are not the culprit

A persistent worry, and one I shared until I tested it properly, is that streaming platforms break sync when they process your stream.

They do not, and the reason is structural.

A stream does not contain audio and video as two separate things that happen to travel together. It contains timing information that says precisely when each piece of audio and each frame of video should be presented, relative to a shared reference. Transcoding, re-encoding your stream into different qualities for different viewers, rewrites the compression while preserving those relative timings. That is the entire contract; a transcoder that broke it would be unusable for everyone.

So if your stream arrives at viewers out of sync, it left your machine out of sync. This is genuinely good news. It means the problem is somewhere you can reach, rather than inside infrastructure you do not control.

The corollary matters too: do not fix sync by adjusting something at the platform. You are correcting a symptom at the far end of a pipe while the leak stays upstream, and the local recording, the copy people will actually watch later, is untouched.

Fixed offset versus drift

These are two different faults with two different causes, and treating one as the other wastes hours.

A fixed offset is constant. Audio leads by the same amount at the start of the service and at the end. This is the normal case, caused by the pipeline asymmetry described above. It is fixed by measuring the offset once and applying it upstream. Once set, it holds.

Drift grows. Sync is fine for the first ten minutes, slightly off at thirty, clearly wrong at sixty. This is a different animal: it means audio and video are being timed by two clocks that do not run at exactly the same rate.

That sounds exotic and is completely mundane. An audio interface has a clock. A camera has a clock. Neither is perfect, and they were never manufactured to agree. A discrepancy far too small to measure directly becomes visible when it accumulates across an hour.

The fix for drift is not a bigger offset... a bigger offset just relocates where in the service it looks correct. The fix is to make both streams answer to a single reference, so that a shared clock governs both and there is nothing to accumulate.

Diagnosing which you have takes one observation. Check sync at the start of your event and again near the end. Same offset means fixed. Growing offset means drift. That single check tells you which problem you own before you touch a setting.

How to actually fix it

  1. Find where audio and video first meet. That is where alignment belongs. Every correction applied after that point is local to one destination.
  2. Measure, do not eyeball. Record something with a sharp, visible transient... a clap works, in frame, close to the microphone. Play it back frame by frame. Count the gap between the frame where hands meet and the sample where the sound starts. That is your offset, as a number.
  3. Apply the delay to audio, upstream, once.
  4. Verify on the recording, not the preview. Preview windows have their own display latency and will lie to you. The recorded file is ground truth.
  5. Check again at the end of a long session. If the offset has grown, you have drift, and you need a shared clock rather than a larger number.

What I learned building this

I spent real time worrying that moving encoding work off the local machine would break sync, and that fear shaped a design decision before I tested it.

When I finally measured, the fear was misplaced. Sync is decided at the encoder, where audio and video are combined. Anything downstream inherits that decision. The thing I was afraid of was not where the risk lived.

There is a general lesson in that, and it is not really about audio. I had a plausible fear, I let it drive a decision, and I had not tested it. Plausible is not the same as true. The measurement took an afternoon; the assumption had been steering me for weeks.

If lip sync is haunting your venue, resist the slider. Find where the two streams meet, measure the gap with a clap, and fix it there. It will hold.

Asked often.

Why is my audio out of sync with my video when streaming?

Audio usually arrives faster than video because video must be captured, converted and encoded while audio does not... the fix is to delay the audio to match, not to speed up the video.

Does streaming to a platform like YouTube cause lip sync problems?

No. Platform transcoding preserves the relative timing it is given, so if your stream arrives out of sync it left your machine that way.

Why does my lip sync drift partway through a service?

Drift that grows over time usually means audio and video are being timed by different clocks, so a small difference in rate accumulates into a visible offset.

I have never done this before. Can I make it worse by trying?

Not permanently. An offset is a number you can set back, and nothing you change here alters recordings that already exist.

Do I have to understand encoding to fix this?

No. Record a clap in frame, count the gap, apply that as an audio delay upstream, and check the recording. The explanation is here if you want it, but the procedure does not require it.

Is a one-button auto-calibration trustworthy for something this fiddly?

It measures a test recording and applies the result, and it refuses a figure outside a plausible range rather than applying it. It is worth verifying its answer on a recording once, the same as any measurement.

What does the software do about drift on its own?

It runs audio and video against a shared reference rather than letting each device's own clock govern its stream, which is what stops a small rate difference accumulating across an hour.

Try it on your Sunday.

Free tier, no credit card. A laptop and the phones in the room.

Start Free →
Why your lip sync drifts (and why it is almost never the camera) | Broadcasteer