Truthring
Understanding the technology · the phenomenon

Deepfake audio

What deepfake audio is

Deepfake audio is synthetic speech made to pass as a recording of a real person: a cloned voice saying words that were never said, or a genuine recording edited so convincingly that the edit is inaudible. Its second effect matters as much as its first, because any real recording can now be denied as fake.

This page is about the phenomenon: what it covers, where it turns up, and the second-order damage that gets far less attention than the first. For what our product does with a file, see deepfake audio detection.


What the term covers, and what it does not

“Deepfake” arrived attached to video and was borrowed for audio afterwards, which is why the word is used loosely. Three things sit under it, and they behave differently enough to be worth separating.

The first is fully generated speech in a cloned voice: a model reads a script in a likeness of someone real. The second is voice conversion, where a person actually performs the words and their timbre is replaced with the target’s, keeping the human timing and emotion underneath. The third is partial manipulation: a genuine recording with a phrase cut out, a name swapped, or a few seconds regenerated in the middle of otherwise authentic audio.

The third is the one that causes the most trouble and receives the least coverage. It is hardest to detect, because the overwhelming majority of the file is real. It is also the most tempting, because a small edit to a real recording is more credible than a whole fabricated one and requires less of the fabricator.

Some things frequently called deepfakes are not. A skilled human impersonation is not one. A recording that has been quoted misleadingly, cut off before a qualifying sentence, or captioned dishonestly is not one either, and it does considerably more damage in practice than synthesis does. Selective editing has always been available and requires no technology at all.


Where audio deepfakes actually show up

Coverage tends to concentrate on the most dramatic possibilities. The distribution of real incidents looks different, and it is worth being specific about categories even where we cannot responsibly name cases.

Fraud is the bulk of it. A cloned voice on a phone call asking for a payment, an approval, or a code. This is unglamorous, high volume, and by a wide margin the most common way an ordinary person encounters synthetic speech. The patterns are catalogued in how voice scams are run.

Impersonation of institutions and public figures. Recorded messages and advertisements using a recognisable voice to lend authority to something the person never endorsed, which is its own well-established pattern.

Disputes between individuals. Employment grievances, family proceedings, school incidents, neighbour conflicts. Audio appears as evidence, and the question of whether it is genuine now has to be asked every time. Being on the receiving end of that is one of the more distressing situations we hear about.

Political and civic contexts. Recorded calls and clips circulated around elections and public controversies. [VERIFY: specific incidents, dates and outcomes with primary sources before citing any of them]

Legitimate uses that look identical. Dubbing, narration, accessibility, and voice preservation for people losing speech to illness all produce synthetic audio of a real person with that person’s agreement. A detector cannot tell consent from theft, which is a limitation of the technology and not a gap to be closed later.


Why audio is the practical medium

Video deepfakes get the attention and audio does the work. The reasons are unglamorous.

There is far less to get right. No face, no lighting, no hands, no lip synchronisation, no continuity across frames. Speech is a one-dimensional signal, and a convincing minute of it demands a small fraction of the effort a convincing minute of video does.

The delivery channels flatter it. Telephone calls, voice notes and forwarded clips are expected to sound compressed, noisy and thin, so the artefacts that would expose a fake are absorbed by a channel everyone already tolerates. A video arriving at the same quality would look suspicious. Audio arriving that way looks normal.

And audio carries social weight that is out of proportion to its evidential strength. A voice feels like presence. People who would scrutinise a photograph will accept a voice note without a second thought, which is exactly why voice notes stopped being proof that a person is real.


The liar’s dividend

Here is the part that gets least attention and may end up mattering most.

Once it is common knowledge that convincing fake audio can be made, everybody caught on a genuine recording acquires a new and entirely reasonable-sounding defence. That is not me. That is AI. A few years ago the denial would have been laughable. Now it is plausible, and plausibility is all a denial needs. Scholars writing about synthetic media named this the liar’s dividend: the advantage handed to bad actors by the existence of forgery, entirely independently of whether anything was forged. [VERIFY: attribution and citation for the coinage of the term]

The dividend is paid automatically. Nobody has to make a fake for it to be collected. It is enough that the audience knows fakes are possible.

What makes it hard to counter is an asymmetry that runs right through this field, and it is the same asymmetry that shapes what a detector can honestly tell you. Establishing that a recording was generated requires finding something. Establishing that a recording was not generated requires ruling everything out, and nothing in a waveform can do that. A clean analysis means no trace of synthesis was found in that file, which is also what a genuine recording that lost its fine detail in a phone codec looks like. So the denial cannot be closed off by analysis alone, and anyone offering to certify a recording as authentic is overselling. We set out where that boundary sits on the limitations page, and how to read a result covers the same asymmetry from the user’s side.


Why the second problem may be the bigger one

Fabrication has a natural ceiling. Each fake has to be made, distributed and believed, and each one causes a bounded amount of harm before it is challenged. Deniability has no such ceiling. It applies at zero marginal cost to every genuine recording that already exists and every one that will be made, and there are vastly more real recordings in the world than fake ones.

The consequences land on the institutions that depend on recorded speech being worth something. A journalist with a leaked call now needs a provenance story as well as the audio. A tribunal weighing a recording has a new dispute to resolve before it can reach the one in front of it. A person who recorded harassment or a threat, as they were advised to, may find that the recording no longer carries the weight they were promised, and that they are now the one being questioned.

That last case is the one that should shape how anyone in this business talks about their product. Overstating what detection can do makes the problem worse rather than better: a certification that cannot be justified will eventually be relied upon and eventually be wrong, and every wrong reliance makes the next honest recording easier to dismiss.


What actually helps

Provenance beats analysis. A recording whose origin, capture device and handling can be accounted for is stronger than one that merely passes a test, and keeping the original file rather than a forwarded copy preserves both the metadata and the fine detail an analysis needs. Getting and keeping a usable recording is the single most useful habit here.

Corroboration beats both. Independent accounts, call records, timings, documents and anything else that has to line up will settle more disputes than a probability ever will.

Where analysis does contribute, it should arrive with its error rates attached and in both directions, which is the discipline this whole site is organised around. A synthetic finding is a claim backed by evidence. A clean result is not clearance. Publishing only the flattering number, as is common, is how a useful tool becomes a liability. If you have a file, you can have it checked; the product page for that work is deepfake audio detection, and the definitional groundwork is in what AI voice detection is.


FAQ

Questions about deepfake audio

What is deepfake audio?

Synthetic or manipulated speech presented as a genuine recording of a specific person. It covers fully generated audio in a cloned voice, real speech converted into someone else’s voice, and authentic recordings with words removed, inserted or regenerated. The common element is the claim that a real person said it.

Is audio easier to fake than video?

Considerably. There is far less to get right, no face or lighting to model, and the usual delivery channels are forgiving: a telephone call or a voice note already sounds compressed and imperfect, so flaws that would betray a video are simply absorbed.

What is the liar’s dividend?

The benefit that accrues to someone caught on a genuine recording once convincing fakes are widely known to exist. The denial no longer sounds absurd, because it might be true. The dividend is paid by the mere possibility of forgery, whether or not any forgery occurred. [VERIFY: attribution and citation for the coinage of the term]

Can a detector prove a recording is real?

No, and this is the hardest limitation to accept. Analysis can find evidence that speech was generated; it cannot produce evidence that speech was not. A clean result is consistent with a genuine recording and equally consistent with a fake whose traces were destroyed by compression on the way to you.

What actually helps against fabricated audio?

Provenance more than analysis. A recording whose origin, device and handling can be accounted for is far stronger than one that merely passes a test, and corroboration from independent sources is stronger still. Detection is useful as one input into that picture, not as a substitute for it.


Reviewed