When only four seconds were generated
Descript is an editing tool. Synthesis inside it is not a way to manufacture a speaker — it is a repair function. Someone recorded two minutes of real speech, fluffed one word, and replaced that word with a generated one rather than book the studio again. The file that results is a real recording with a synthetic patch in it. That is the hardest thing on this site to give an honest verdict on.
Measured on [VERIFY: n] files built by inserting generated words into genuine recordings from [VERIFY: n] speakers, on [VERIFY: date], held out of training. Composition: benchmark.
Why a mixed file breaks the usual logic
Every other page on this site rests on a question with two answers: was this audio captured or manufactured? A repaired recording answers both, and the proportions are lopsided. If four seconds of a two-minute file were generated, then roughly ninety-seven per cent of the evidence in that file is genuine evidence of a real microphone in a real room, and it is genuine because a real microphone in a real room produced it.
A whole-file verdict is an average over that evidence. The average is dominated by the majority, which is human, and it is correct about the majority. It is simply answering a question you did not ask. You wanted to know whether any part was generated; the number you got describes the file as a whole.
Worse, the patch is engineered to blend. A repair tool has an advantage no impersonation tool has: it can condition the generated fragment on the surrounding real audio — the speaker’s own voice, the room, the noise floor, the pace of the sentence it is being dropped into. The target is not sound like a person
but sound like this recording
, which is a narrower and much more achievable target. The seams that a fully generated clip shows at every sentence boundary appear here at exactly two boundaries, and both are inside a word.
What we can and cannot say
| What you have | What a verdict can support | What it cannot |
|---|---|---|
| Whole file, one submission | Whether the file as a whole reads generated | Whether a short insert exists inside it |
| Short consecutive segments | Which region reads differently from its neighbours | The exact word boundary of an edit |
| A file with a suspected cut | Whether the acoustics change across that point | Whether the change was a synthesis or a splice of other real audio |
| Compressed export of an edited file | Very little; the insert is usually the first thing lost | Any confident statement about the edit at all |
How to actually look for an insert
The method is unglamorous and it works better than any single verdict.
Cut the file into short consecutive windows
A few seconds each, overlapping slightly, in order, with no gaps. You are not trying to find the suspicious bit — you are building a profile across the whole recording so that the suspicious bit has something to stand out against.
Read the sequence, not the numbers
Absolute values matter less than shape. Twenty windows sitting in one range and one window sitting well outside it is the signal. Twenty windows scattered everywhere means the recording is too noisy for this to work at all.
Listen to the outlier
Once a window is flagged, a person should hear it in context. Repairs often coincide with something audible — a breath that is missing, a room tone that steps, a word whose consonants sit slightly forward of the ones around them. The tool narrows where to listen; it does not replace listening.
The honest limit: below a certain insert length there is not enough material in any window to reach a confident reading, and the edit is effectively invisible to us. That length is [VERIFY: n] seconds on clean audio and longer on anything compressed.
The asymmetry is sharper here than anywhere else. Likely synthetic requires positive evidence, so on a mixed file it is a strong result — something in that audio was manufactured. Likely human is much weaker: on an edited file it can mean the genuine majority swamped a real insert, or that an export destroyed the evidence. A clean human verdict on a two-minute file is not a certificate that nothing was changed.
The part detection cannot decide
Suppose the analysis works perfectly and tells you that seconds forty-one to forty-five of a recording were generated. You still do not know whether anything wrong happened. A presenter fixing a mispronounced surname, a producer removing a cough, and someone altering what a person is heard to have agreed to all leave the same kind of trace. The signal is identical; the meaning is entirely in the context.
This matters most in the settings where recordings are used as evidence — a disciplinary hearing, an insurance claim, a dispute over what was said on a call. There, the useful output is not this is fake but this region does not match the rest of the recording, and someone should explain why. That is a much smaller claim, and it is one a recording can actually support.
Where money, employment, a legal claim or someone’s safety turns on the answer, treat a segment reading as a prompt to ask for the original project file or the source recording. An edit that was innocent is usually easy for its maker to account for.
Questions
Can you detect a single replaced word?
Sometimes, on clean audio, using segment-by-segment checking rather than a single submission. Below [VERIFY: n] seconds of generated material we usually cannot, and we would rather say so than produce a number that sounds decisive.
My whole-file check said likely human. Is the recording clean?
It means the file as a whole read that way, which is the expected outcome for a mostly genuine recording even when an insert is present. If your question is about a specific moment, check that moment on its own.
Is editing a recording with synthesis dishonest?
Not inherently, and most of it is routine production work. The distinction that matters is whether the edit changes what a listener would understand the speaker to have said. Detection can point at where an edit is; only the surrounding facts say what it did.
Can you tell an AI repair from a plain cut-and-splice?
Not reliably. Both produce a discontinuity in the recording’s acoustics. A generated patch has additional characteristics we look for, but a well-made splice of the speaker’s own real audio can read similarly, and we do not distinguish them with confidence.
What should I submit if the file is long?
Consecutive short windows across the whole file, kept in order and named so you can reconstruct the sequence. Send the original export rather than a version that has been through a messaging app — on a mixed file, compression takes the insert first.
Signature last retested [VERIFY: date] against Descript model version [VERIFY: verify]. Partial-synthesis rates are re-measured monthly and depend heavily on insert length.
Reviewed