Truthring
Coverage · the hardest case

When only four seconds were generated

Descript is an editing tool. Synthesis inside it is not a way to manufacture a speaker — it is a repair function. Someone recorded two minutes of real speech, fluffed one word, and replaced that word with a generated one rather than book the studio again. The file that results is a real recording with a synthetic patch in it. That is the hardest thing on this site to give an honest verdict on.

Partial synthesis Signature held since [VERIFY: date] · last retested [VERIFY: date]
Fully generated clip[VERIFY]%
Insert of [VERIFY: n] sec in a [VERIFY: n] sec file[VERIFY]%
Same insert, checked segment by segment[VERIFY]%
Unedited human recording wrongly flagged[VERIFY]%

Measured on [VERIFY: n] files built by inserting generated words into genuine recordings from [VERIFY: n] speakers, on [VERIFY: date], held out of training. Composition: benchmark.


Why a mixed file breaks the usual logic

Every other page on this site rests on a question with two answers: was this audio captured or manufactured? A repaired recording answers both, and the proportions are lopsided. If four seconds of a two-minute file were generated, then roughly ninety-seven per cent of the evidence in that file is genuine evidence of a real microphone in a real room, and it is genuine because a real microphone in a real room produced it.

A whole-file verdict is an average over that evidence. The average is dominated by the majority, which is human, and it is correct about the majority. It is simply answering a question you did not ask. You wanted to know whether any part was generated; the number you got describes the file as a whole.

Worse, the patch is engineered to blend. A repair tool has an advantage no impersonation tool has: it can condition the generated fragment on the surrounding real audio — the speaker’s own voice, the room, the noise floor, the pace of the sentence it is being dropped into. The target is not sound like a person but sound like this recording, which is a narrower and much more achievable target. The seams that a fully generated clip shows at every sentence boundary appear here at exactly two boundaries, and both are inside a word.

What we can and cannot say

What you haveWhat a verdict can supportWhat it cannot
Whole file, one submissionWhether the file as a whole reads generatedWhether a short insert exists inside it
Short consecutive segmentsWhich region reads differently from its neighboursThe exact word boundary of an edit
A file with a suspected cutWhether the acoustics change across that pointWhether the change was a synthesis or a splice of other real audio
Compressed export of an edited fileVery little; the insert is usually the first thing lostAny confident statement about the edit at all

How to actually look for an insert

The method is unglamorous and it works better than any single verdict.

1

Cut the file into short consecutive windows

A few seconds each, overlapping slightly, in order, with no gaps. You are not trying to find the suspicious bit — you are building a profile across the whole recording so that the suspicious bit has something to stand out against.

Segment length: [VERIFY: seconds]

2

Read the sequence, not the numbers

Absolute values matter less than shape. Twenty windows sitting in one range and one window sitting well outside it is the signal. Twenty windows scattered everywhere means the recording is too noisy for this to work at all.

Interpretation: relative, not absolute

3

Listen to the outlier

Once a window is flagged, a person should hear it in context. Repairs often coincide with something audible — a breath that is missing, a room tone that steps, a word whose consonants sit slightly forward of the ones around them. The tool narrows where to listen; it does not replace listening.

Confirmation: human review

The honest limit: below a certain insert length there is not enough material in any window to reach a confident reading, and the edit is effectively invisible to us. That length is [VERIFY: n] seconds on clean audio and longer on anything compressed.

The asymmetry is sharper here than anywhere else. Likely synthetic requires positive evidence, so on a mixed file it is a strong result — something in that audio was manufactured. Likely human is much weaker: on an edited file it can mean the genuine majority swamped a real insert, or that an export destroyed the evidence. A clean human verdict on a two-minute file is not a certificate that nothing was changed.


The part detection cannot decide

Suppose the analysis works perfectly and tells you that seconds forty-one to forty-five of a recording were generated. You still do not know whether anything wrong happened. A presenter fixing a mispronounced surname, a producer removing a cough, and someone altering what a person is heard to have agreed to all leave the same kind of trace. The signal is identical; the meaning is entirely in the context.

This matters most in the settings where recordings are used as evidence — a disciplinary hearing, an insurance claim, a dispute over what was said on a call. There, the useful output is not this is fake but this region does not match the rest of the recording, and someone should explain why. That is a much smaller claim, and it is one a recording can actually support.

Where money, employment, a legal claim or someone’s safety turns on the answer, treat a segment reading as a prompt to ask for the original project file or the source recording. An edit that was innocent is usually easy for its maker to account for.

Questions

Can you detect a single replaced word?

Sometimes, on clean audio, using segment-by-segment checking rather than a single submission. Below [VERIFY: n] seconds of generated material we usually cannot, and we would rather say so than produce a number that sounds decisive.

My whole-file check said likely human. Is the recording clean?

It means the file as a whole read that way, which is the expected outcome for a mostly genuine recording even when an insert is present. If your question is about a specific moment, check that moment on its own.

Is editing a recording with synthesis dishonest?

Not inherently, and most of it is routine production work. The distinction that matters is whether the edit changes what a listener would understand the speaker to have said. Detection can point at where an edit is; only the surrounding facts say what it did.

Can you tell an AI repair from a plain cut-and-splice?

Not reliably. Both produce a discontinuity in the recording’s acoustics. A generated patch has additional characteristics we look for, but a well-made splice of the speaker’s own real audio can read similarly, and we do not distinguish them with confidence.

What should I submit if the file is long?

Consecutive short windows across the whole file, kept in order and named so you can reconstruct the sequence. Send the original export rather than a version that has been through a messaging app — on a mixed file, compression takes the insert first.

Signature last retested [VERIFY: date] against Descript model version [VERIFY: verify]. Partial-synthesis rates are re-measured monthly and depend heavily on insert length.

Reviewed