Podcast and Spoken-Word Audio
How to record a double-ender so the two tracks do not drift out of sync
Published October 2, 2026
Drift is not a mistake you made lining the files up. It is two computers recording at slightly different real speeds, and you cannot remove it without a shared clock, which two people in two cities do not have. So the job is to make drift small, measurable and correctable. Before the interview: agree one sample rate at both ends, have each person record locally from one wired input, have the host also record the call itself as a reference, and mark the end of the session as well as the start. Then run a sixty-second test. A 100 ppm clock difference separates two recordings by about 360 ms over an hour, and consumer devices have measured far worse than that.
A double-ender is the standard way to record a remote interview properly: instead of keeping the compressed call audio, each person records themselves locally and you combine the two clean files afterwards. It solves the quality problem completely. It introduces a quieter one that people usually discover halfway through the edit, when the clap at the top lines up perfectly and the laugh forty minutes in lands a visible distance late.
That is drift, and it is predictable rather than mysterious. This article is about preventing it at the setup stage and knowing how much of it you have. If you already have two files that have drifted apart and the recording is over, the repair is a different job, covered in how to fix audio that is out of sync in REAPER.
Two different problems get called the same thing
Before anything else, work out which one you have, because the fixes have nothing in common.
- A nominal sample rate mismatch. One end recorded at one rate and the file is being played at another, so it is wrong by a large, constant proportion from the very first second, and the pitch is obviously off. That is not drift, and it has its own fix in how to fix audio recorded at the wrong sample rate.
- A clock accuracy difference. Both ends genuinely recorded at the rate they were set to, but the crystal oscillator in each machine does not run at exactly the nominal frequency. One device's real rate might be a fraction above 48 kHz and the other's a fraction below. The files start together and separate slowly. That is drift, and it is what the rest of this article is about.
The ASIO4ALL documentation puts the mechanism plainly: whenever two clocks are not identical, devices will slowly drift apart, and a 100 ppm difference means two recordings will diverge by 360 ms every hour.
Clock accuracy is cheap to achieve, and the pitch error involved is tiny. Dan Lavry's paper on clock accuracy notes that a shift of one cent, which is under 500 ppm, is below the limit of what people can hear. So nobody notices that one track is very slightly sharp. What they notice is that the reaction lands late.
The same paper makes the point that decides how you should think about the whole thing: holding two independently clocked tracks within one microsecond of each other over three minutes would require a total error inside 0.0056 ppm. No consumer setup does that. Zero drift is not a setting you can switch on. Everything below is about making it small enough to ignore, or large but known.
How much drift to expect
The arithmetic is simple. A clock difference expressed in parts per million multiplied by the elapsed time gives you the separation.
| Clock difference | Separation per minute | Separation after one hour |
|---|---|---|
10 ppm | 0.6 ms | 36 ms |
50 ppm | 3 ms | 180 ms |
100 ppm | 6 ms | 360 ms |
500 ppm | 30 ms | 1.8 s |
1000 ppm | 60 ms | 3.6 s |
Those higher rows are not hypothetical. A peer-reviewed study of time drift in hand-held recording devices, presented at MultiMedia Modeling in 2015, measured oscillator error across sixteen consumer devices and found misalignment reaching about 60 ms per minute in the worst case, which is the bottom row of that table. The same study found that drift is not even constant: devices needed roughly ten minutes to reach working temperature, and the rate moved with temperature.
The devices measured were phones, tablets, media players and camcorders, not audio interfaces. Those figures tell you what consumer hardware can do, which matters because half of all double-enders are recorded on a phone or a laptop. They are not a prediction for your own rig. The only number that applies to your gear is the one you measure, and the test below is how you get it.
The setup that keeps drift small and fixable
Agree one sample rate and write it down
Which rate you pick matters far less than both ends using the same one. Send the guest the number in writing rather than assuming a default. In REAPER, the project rate lives in File › Project Settings and the hardware rate in Options › Preferences › Audio › Device. Matching them does not stop drift, but it removes the much larger, much more obvious error from the picture so that what is left is only drift.
One wired input per machine, recorded locally
Record from the interface or USB microphone that the recording computer is clocked to, and record to a local file rather than keeping the call stream. Every extra device in the chain that runs on its own clock is another place where two rates have to be reconciled, which is the same mechanism as the drift you are trying to limit.
Have the host record the call as a reference
The call recording is poor quality and you will throw it away, but it is continuous and it contains both voices, which makes it the ruler you align the two good files against. This is the approach the free Double Ender utility is built on: local recordings are synced against a reference file so they start together. Record it every time, even when you are sure you will not need it.
Mark the end, not just the start
Everyone knows about the clap at the top. Almost nobody claps at the end, and the tail marker is the single most useful thing in this list. With only a head marker you can see that something went wrong. With a head and a tail marker you can calculate the rate and correct it. Use a count-in, a clap or a short tone, done by both people, captured on both local files.
Record one continuous file
Do not pause and resume. A pause puts a second, unrelated alignment problem on top of the drift, and it breaks the arithmetic that the tail marker exists to make possible. If you must stop, treat what follows as a separate recording with its own head and tail markers.
Run a sixty-second test before the real thing
Clap, talk for a minute, clap again, then both of you send the file. Measure it with the method below. One minute of testing tells you whether this pairing of machines drifts enough to matter, and it is the only step here that produces a real number for your actual setup.
How to measure the drift you actually have
Line the two files up on their head markers. Then look at the tail markers and note the gap between them. Divide that gap by the elapsed time between the markers.
A worked example. The head markers are aligned, and after 58 minutes of recording the tail markers sit 420 ms apart. That is 0.420 seconds over 3,480 seconds, which is about 121 ppm, or roughly 7 ms per minute. Now you know the rate rather than the symptom, and you can decide whether it needs correcting.
What to do with the number is a judgement rather than a rule, so here is the reasoning instead of a threshold. Conversational timing is much more forgiving than lip sync: a reaction that is a few tens of milliseconds early or late reads as normal human overlap. A reaction several hundred milliseconds out reads as a bad edit. So a total separation in the low tens of milliseconds across the whole episode usually needs nothing but the head alignment, and a separation of hundreds of milliseconds or more needs the file stretched, which is the method in the sync repair article.
Because drift is not perfectly linear, applying one measured rate as a single correction leaves some residual error, and the study above suggests the worst of it sits in the first ten minutes while the hardware is still warming up. If you need it tight, measure the gap at two or three points through the episode instead of only at the end, and correct in sections.
What this will not do
- It will not make drift zero. Only a shared clock does that, and a remote interview by definition does not have one.
- It will not repair a recording you have already made. Nothing here is retroactive. Two files already in hand need the repair method, not this checklist.
- It will not fix a nominal rate mismatch. That is a bigger, constant error with a different cause and a different fix.
- It will not improve how the guest sounds. Drift and quality are unrelated. A thin, noisy or echoey guest track is a separate cleanup job.
- It will not fix dropouts or missing audio. A file that lost samples is incomplete, not drifting, and stretching it by a constant amount makes it worse rather than better.
- It does nothing if you only ever had the call recording. One recording means one clock and no drift, and your problem is quality instead.
If it still will not line up
- You aligned on speech instead of the marker. Align on the clap or tone. The first word is not a reliable reference because the two people did not start talking at the same instant.
- The gap is constant rather than growing. Compare the head gap with the tail gap. If they are the same, this is a fixed offset, which is much easier: slide one file and stop.
- The lengths differ by a suspiciously round proportion. If one file is about 1.0884 times the length of the other, that is the 48 kHz against 44.1 kHz ratio, so it is a rate mismatch rather than drift.
- The gap jumps rather than grows. Look for a discontinuity. Something paused, dropped out, or was stitched together, and no single stretch will fix a step change.
- You are judging the final tightness on the call recording. Use the reference to find the rough alignment, then make the final call on the two local files, because those are the ones that get published.
Build it into the session, not into the edit
All of the above costs about three minutes of setup, and most of it is a note you can send a guest in advance alongside the rest of your podcast recording setup. The sample rate, the local recording, the reference, the head clap and the tail clap. Once the two files are aligned and level-matched, the next job is making the two voices sit together, which is host and guest loudness matching.
Drift comes from two clocks that are not identical, so it cannot be eliminated, only made small and measurable. Agree one sample rate, record locally from one wired input at each end, keep the call as a reference, and mark the end of the session as well as the start. Measure the gap between the tail markers, divide by the elapsed time, and you have a number you can act on instead of a sync problem you can only see.
Sources
- ASIO4ALL, “Device aggregation” - states that whenever two clocks are not identical devices will slowly drift apart, and that a 100 PPM difference means two recordings will diverge by 360 ms every hour. asio4all.org/device-aggregation
- M. Guggenberger, M. Lux and L. Böszörmenyi, “An Analysis of Time Drift in Hand-Held Recording Devices”, MultiMedia Modeling (MMM 2015), Lecture Notes in Computer Science vol. 8935, Springer, pp. 203-213 - measured oscillator error across consumer hand-held devices, with temporal misalignment reaching about 60 ms per minute in the worst case, and reported a warm-up period before devices reach working temperature. doi.org/10.1007/978-3-319-14445-0_18
- D. Lavry, Lavry Engineering, “Clock Jitter and Clock Accuracy for Digital Audio” - distinguishes clock accuracy from jitter, notes that a shift of one cent (under 500 ppm) is below the audible limit, and that restricting drift between two tracks to within one microsecond over three minutes would require total error within 0.0056 ppm. lavryengineering.com
- The Incomparable, “Double Ender” - a free utility that syncs local recordings to a reference file; states that different audio devices can record at slightly different speeds, causing audio to drift out of sync over time. theincomparable.com/doubleender
- Cockos, “REAPER Quick Start Guide” - the audio device settings are reached via
Options › Preferences › Audio › Device, and item properties via right-click on an item orF2. reaper.fm/guides - Cockos, “Up and Running: A REAPER User Guide” (troubleshooting guide) - the project sample rate is set in
File, Project Settings. reaper.fm/guides/ReaperTroubleshootingGuide.pdf - Cockos, REAPER User Guide (current version 7.81, released September 28, 2026). reaper.fm/userguide