Voice conversion
A human being says every word; a model replaces only who it sounds like. This is the attack that can answer your questions, and it is routinely reported as cloning, which it is not.
Voice conversion takes speech a real person is actually performing and re-renders it in a different voice, preserving the original words, timing, emphasis and emotion. A live human drives the delivery; the model replaces only the vocal identity. It can run in real time, during a call, on whatever the speaker says next.
Why it is the hard case for detection
Detection works by finding what a generator had to invent and could not invent perfectly. Conversion narrows that surface dramatically, because most of the recording is not invented at all.
The rhythm is human, because a person produced it. The hesitations, the false starts, the breath before a long sentence, the way the pitch falls when the speaker is tired — all human. The room is a real room. The microphone is a real microphone, with its own noise floor and handling sounds. What the model supplies is the spectral character of a different larynx and vocal tract laid over that performance.
So the evidence available to a detector is confined to a narrower band of the signal than it is with fully generated speech, and it competes with everything genuine around it. Then the telephone arrives and removes a further share of exactly that band. Our note on phone and codec audio explains why a call is the worst place to look for the traces that remain, and it is precisely where conversion is used.
It breaks the advice most people have been given
Nearly every consumer guide to voice fraud rests on the assumption that the caller is a recording. Ask something only the real person would know. Ask an unexpected question. Interrupt and see whether they respond.
Against a converted voice, none of that holds. There is a person on the line. They hear the question, think about it, and answer in their own words — arriving in a voice that is not theirs. They can be surprised, they can laugh, they can be annoyed at being doubted, because a human is genuinely being surprised, laughing and getting annoyed.
What survives is the procedural defence, and it survives untouched: end the call and dial back on a number you already had. The person running a converted voice cannot answer the handset belonging to the person they are impersonating. That is the reason our guide to hearing the difference spends most of its length arguing against listening.
A concrete case
A company hires remotely. A candidate interviews over video with the camera off, citing bandwidth, and performs strongly. The voice on the call belongs to a person who is not the person on the CV, converted in real time to match the short introduction video the real candidate had recorded earlier.
Nothing in the interview behaves like a recording. Answers are specific, follow-up questions land, jokes work. The failure was structural rather than acoustic: the identity was never verified through a channel independent of the call. Voice fraud in hiring covers how that check is built into a process, and what recruiters can reasonably do about it.
Commonly confused with: voice cloning
Reporting tends to call every impersonation a clone. The distinction is simply where the performance comes from.
Cloning generates the whole utterance from text. No human is speaking, and the result is a fixed clip — strong for voicemails and voice notes, weak in conversation, because it cannot answer anything it was not written to say.
Conversion transforms a live performance. It is weak at scale, because it costs an operator’s time for every call, and strong in exactly the situations where a decision is being negotiated in real time. If you are designing a control, cloning is the volume threat and conversion is the targeted one.
Where Truthring stands on it
Conversion is the category most likely to return unclear, and we would rather return that than a confident answer we cannot support. Unclear means the clip did not carry enough usable signal to decide — and on a converted, telephone-quality call it frequently does not.
It also sharpens the asymmetry between our two other verdicts. A likely synthetic result on converted speech is meaningful, because something positively matched. A likely human result is worth very little there, since a human really was speaking and the one machine-made component may not have survived the codec. We list this as a known weak point on the limitations page rather than in a footnote, and the research notes track what we are doing about it. [VERIFY: publish a per-category breakdown showing conversion separately from full synthesis once the evaluation set covers it]
Questions this term raises
Is voice conversion the same as voice cloning?
No. Cloning generates a whole utterance from text with nobody speaking. Conversion transforms speech a live person is performing, keeping their timing and delivery and replacing only the vocal identity. The practical difference is that a converted voice can hold a conversation.
Can a detector catch a converted voice?
Sometimes, and less reliably than it catches fully generated speech. Most of the recording is genuinely human, so there is less for the analysis to find, and a phone codec attacks the part of the signal where the machine-made component sits.
Does asking a personal question defeat it?
Not on its own. A real person is listening and can answer, guess, or deflect. Questions test knowledge, which an operator may have researched. Calling back on a number you already hold tests control of the line, which they do not have.
Related terms
Reviewed