Voice clone detection, and the three things people call cloning
A voice clone detector estimates whether recorded speech was produced by a model imitating a specific person rather than spoken by that person. Truthring reports a likelihood together with a published error rate, names the generating system where it can, and makes no claim about whose voice was copied.
Three different techniques get filed under one word, and they leave different amounts of evidence behind. Knowing which one you are dealing with is most of the reason a result comes back strong or comes back barely worth having.
Three techniques, one label
Text to speech
Text goes in, speech comes out, in a stock or designed voice that belongs to nobody. Nothing about it targets a real person. Detection is comparatively tractable: the whole performance was invented, so nothing in the timing or the breathing came from a human being.
Voice cloning
A sample of one person’s speech conditions the model, which then reads any script in a likeness of that voice. This is what appears in fraud, and what people mean when they say a voice was cloned. Still fully generated, so still detectable in principle.
Voice conversion
A human performs, live or recorded, and the model rewrites the timbre so it sounds like a different person. The words, the pauses, the laugh, the hesitation are all real. Only the identity was swapped. The hardest of the three by a wide margin.
Why voice conversion is the difficult case
A useful share of what exposes generated speech is behavioural rather than acoustic. Synthesis systems learn from clean, well-lit, read material, and they are least convincing at the places where recorded conversation is untidy: a word restarted mid-syllable, a breath taken in the wrong place, the way a sentence collapses when the speaker changes their mind about it. Those are the moments a detector leans on.
Voice conversion removes that advantage in one step. A person really did restart the word and really did change their mind, and the model preserved all of it while altering only the colour of the voice. What is left to find is thinner — artefacts around the conversion itself, inconsistencies between the claimed vocal apparatus and the way it resonates, seams where the model handled an unusual sound badly. Real evidence, but less of it, and the first thing a compressed channel destroys.
The practical implication for anyone assessing a suspicious recording is uncomfortable and worth stating plainly. A confident likely human on a clip you suspect was converted deserves less trust than the same verdict on a clip you suspect was read out by a text-to-speech system. The floor is lower here, and we would rather you knew where the floor was.
What makes cloning a named person its own problem
Detecting generated speech in the abstract is a signal-processing question. Detecting a clone of someone who exists brings in three complications that have nothing to do with signal processing.
The source material is usually public, so nothing was breached
Conference recordings, podcast appearances, earnings calls, videos posted to a family account, a voicemail sitting on a colleague’s phone. No system was compromised to obtain any of it. There is no incident to investigate and no moment at which the person could have been warned.
The target is chosen because they are recognisable
Which means the clip usually arrives with a plausible story attached and lands on someone primed to believe it. Detection sits at the end of that sequence, after the recipient has already decided the voice sounds right. Voice fraud rarely fails on audio quality; it fails when someone rings back on a number they already had.
The consequences of being wrong are asymmetric
Missing a clone may cost money. Wrongly flagging a genuine recording as a clone accuses a real person of something they did not do, in a dispute where the recording was probably their defence. This is why the rate at which authentic speech gets flagged is published beside the detection rate on the accuracy page, rather than left out of the marketing.
What a clone verdict is a claim about
It is a claim about a recording. It is not a claim about a person’s conduct, and the distance between those two sentences is where most of the harm in this field is done.
Truthring examines a file and reports the likelihood that the speech in it was manufactured. It does not know who uploaded the file, who first published it, who is depicted, or what anyone intended. A file arrives having already passed through hands we cannot see. Somebody forwarding a synthetic clip is far more often a victim of it than its author.
So the honest form of a result is narrow. This audio shows evidence consistent with synthesis, at this confidence, under these conditions, with this false positive rate. Not this person faked a recording. Every report carries a reference code and a model version precisely so the narrow claim can be checked by someone who disagrees with it, which is what separates a piece of evidence from an assertion.
Before you analyse anything, do the cheap thing. If a voice you know has asked you for money, access or urgency, end the call and ring the number you already had stored. A cloned voice cannot answer a phone you dialled. Nearly every voice fraud that works, works because the target stayed inside the channel the attacker picked.
Systems that clone a specific voice
Coverage notes for the commercial cloning services, with what each one is used for and how its output behaves under analysis: ElevenLabs, Resemble AI, PlayHT, Descript Overdub, Lovo, Replica Studios, Camb.ai and Fish Audio.
Open-weight cloning is a separate category, because a model running on a laptop leaves no account, no invoice and no vendor-side record of what was generated: Coqui XTTS, Tortoise TTS, Bark, VALL-E, and the category overview.
The complete list of systems we hold signatures for is on the detector page.
Questions
What is the difference between voice cloning and text to speech?
Text to speech invents a voice, or uses a stock one, and reads whatever you type. Voice cloning targets a particular person and rebuilds their voice from recordings of them. The machinery overlaps; the consequence does not. Only one of them can be used to impersonate someone who exists.
What is voice conversion, and why is it harder to detect?
In voice conversion a real person speaks and the model rewrites their voice to sound like someone else, keeping the original timing, breathing and emotion. Because a human performance is underneath, the behavioural cues that give text to speech away are genuine. Only the timbre was manufactured, which leaves less to find.
Can Truthring tell me whose voice was cloned?
No. Truthring does not identify speakers and holds no register of voiceprints. It can say that speech in a recording appears to have been generated, and sometimes which system generated it. Matching a voice to a named person is speaker identification, a different discipline with a different and worse error profile.
Does a clone verdict prove someone was impersonated?
It does not. A verdict is a statement about one audio file, and files travel. The person who sent it may have made it, forwarded it, or recorded it off a screen. Treating a synthetic result as a finding about a named individual’s conduct is a leap the analysis does not support.
How much recorded speech does a usable clone need?
Less than most people assume, and the threshold keeps falling. Anyone who has spoken on a podcast, given a conference talk, posted a video or left a long voicemail should treat their voice as reachable material rather than as something private.
Method and error
Related detection
If it has happened to you
Reviewed