How voice cloning works
Voice cloning is the use of a short recording of someone speaking to condition a speech synthesis model on that person’s vocal identity, so the model can then read any text in a recognisable likeness of their voice. Roughly a minute of clear speech is usually enough for a likeness that survives a phone call.
This explains the shape of the capability and what it costs an attacker, which is what a journalist, a security lead or a worried parent needs. It is not a guide to doing it, and it deliberately omits the steps.
What a model takes from a recording
A speech synthesis model has two jobs that can be separated. One is producing speech at all: turning text into a sequence of sounds with plausible timing, stress and intonation. The other is producing speech that sounds like a particular person.
Cloning is the second job. From a sample of someone talking, a system derives a compact representation of what makes that voice recognisable, and this representation is then used to steer generation. The output is not a rearrangement of the recording. Nothing from the sample is being played back. The model is generating new speech that has been pointed at a target identity, which is why a clone can say sentences the person never uttered and never would.
Two broad approaches exist and they have different economics. One conditions a general model on a short reference clip at the moment of generation, needing no training run and very little audio. The other adapts a model to a specific speaker using more material and produces a closer likeness. The first is fast and cheap; the second is better. Fraud only needs the first.
A third mode matters for live calls: converting one person’s speech into another’s voice as they talk, preserving their timing and emotion while replacing their timbre. This is why an operator can improvise, answer unexpected questions and react to a hesitant listener while sounding like someone you know. It is also, for reasons covered in how detection works, harder to detect than fully generated audio, because the performance underneath is genuinely human.
Why roughly a minute is enough
The figure surprises people, and understanding why it is so low is more useful than memorising it.
A minute of ordinary connected speech already contains most of what a model needs: the pitch range a person habitually uses, the resonance of their particular vocal tract, their consonant sharpness, their rhythm, most of the phonemes of their language. It is a small sample statistically and a rich one perceptually.
The second reason is that the target is lower than it looks. A clone does not have to survive a mastering engineer with studio monitors. It has to survive a mobile phone, which discards most of the frequency range, compresses hard, and hands the listener a signal that already sounds nothing like the person in the room. Half the difference between a good clone and a great one is thrown away by the channel before anyone hears it.
The third reason is human. Recognition of a familiar voice is fast, automatic and generous, and it gets more generous under stress. A listener told that the speaker is crying, injured or frightened will explain away every remaining discrepancy without noticing they have done it. That is the mechanism the family emergency call is built on, and it is not a failure of attention.
Where the source material comes from
Almost always from something the person published themselves, or agreed to have published.
A podcast episode is an hour of clean, close-miked, single-speaker audio. A recorded conference talk is the same. So are webinars, lecture captures, product demos, investor calls posted for transparency, council meetings streamed as a public duty, radio interviews, and the growing archive of short vertical video where people talk straight into a phone. An outgoing voicemail greeting is a few seconds of unusually clean speech that anyone in the world can collect by ringing a number that does not get answered.
This is why the executives most exposed are the ones whose jobs require them to be heard. Voice fraud aimed at finance teams works partly because a chief executive’s voice is a marketing asset that has been distributed deliberately. Nobody is going to stop giving talks, and telling them to would be silly advice.
Two practical reductions are still worth making. A short, impersonal voicemail greeting gives away less than a long chatty one. And a public post announcing that someone is travelling, in hospital, or unreachable supplies the detail that makes a story fit, which is a different kind of leak and often the more useful one to an attacker.
Consent is a checkbox, not a check
Legitimate cloning is a real business. Audiobook narration, dubbing into languages a performer does not speak, accessibility for people losing their voice to illness, and the preservation of a voice before surgery are all uses that ordinary people would recognise as good.
The usual control on all of it is an attestation. When an account holder uploads a sample, they confirm that the voice is theirs or that they have the rights to it, and that confirmation is generally where the checking ends. The point to grasp is what an attestation is for. It allocates liability. It does not establish a fact. No register exists against which an uploaded voice could be matched to a person who agreed, and building one would create a biometric database with problems of its own.
Some services layer additional friction on top, and the details differ by vendor and change often. [VERIFY: current consent and verification controls at each named provider before publishing specifics]
The consequence for anyone whose voice has been used without permission is uncomfortable and should not be dressed up. Recourse runs through platform reporting, the terms of a specific service and whatever local law covers likeness and impersonation, and it is slow. What can realistically be done sets out the honest list.
What this page will not tell you
There is no walkthrough here, no comparison of which system produces the most convincing likeness, and no discussion of settings. That omission is deliberate and it is a standing policy rather than a decision about this article. Our acceptable use position is that describing a capability is legitimate and lowering the cost of exercising it is not.
The same rule shapes what we publish about detection. We describe why the analysis works and where it fails, and we do not publish material whose main value would be to help someone strip a signature out of a file. That leaves some questions on this site permanently unanswered, and we would rather carry that than the alternative.
If you have a recording and want to know whether it was generated, the check page is the place to start, and reading your result explains why a clean answer is weaker than a flag.
Questions about cloning
How much audio does it take to clone a voice?
Less than most people expect. Roughly a minute of clear speech is usually enough to produce a likeness that survives a phone call, and a phone call is a forgiving channel. More audio buys a better likeness rather than a possible one, so the practical threshold has already been crossed for anyone who has spoken in public.
Where do people get the sample?
From material that was published on purpose. Podcast appearances, conference talks, webinars, lecture recordings, product demos, social video, livestreams, and outgoing voicemail greetings. Nothing has to be stolen. A single interview usually contains more clean speech than is needed.
Does a cloned voice sound identical to the real person?
Close enough for the situations that matter. It reproduces timbre and delivery style, not knowledge or history, and on a compressed phone line the remaining differences are largely gone. Distress in the script also covers imperfections, because people accept that a frightened relative does not sound like themselves.
Do the companies that offer cloning check consent?
Typically they ask the account holder to attest that they hold the rights to the voice. An attestation is a contractual position, not a verification: nobody is comparing the uploaded sample against a register of people who agreed. Some services add further checks. [VERIFY: current consent controls at named vendors before citing any of them]
Can I stop my voice being cloned?
Not reliably, and it is worth saying so plainly rather than offering comfort. You can shorten a voicemail greeting and reduce what is published, but you cannot withdraw a talk that is already online. The defence is procedural: an agreed passphrase and a callback rule, which work regardless of how good the clone is.
When a voice is misused
Protective measures
Reviewed