How to tell if a voice is AI
You usually cannot tell by listening, and trying to is the wrong move. Hang up and call back on a number you already have. That single procedure beats every listening cue on this page.
Why the honest answer is the useful one. Every guide that hands you a list of tells is training you to stay on the call and deliberate, which is exactly what the person on the other end needs. Deciding takes seconds; verifying takes a minute and cannot be argued with.
The cues that used to work and no longer do
Most advice still circulating was written against an earlier generation of speech synthesis. Those systems really did have audible fingerprints. Current ones largely do not, and the ones that remain are unreliable enough that acting on them is worse than not listening at all.
- Flat or robotic intonation. Older systems produced a monotone that human ears caught immediately. Modern generators model prosody directly and can carry stress, hesitation and emotional colour convincingly.
- Wrong emphasis on the wrong word. The classic tell — stress landing on a preposition, a question ending flat. Now rare enough in normal speech that its absence proves nothing and its presence is more likely a nervous human.
- No breathing. Breath sounds are routinely synthesised, and in any case a phone codec strips them so thoroughly that a real speaker often sounds breathless too.
- Perfectly even pacing. Deliberate irregularity is inserted by design in most current systems. A caller reading from a script will often sound more mechanical than a clone will.
- Digital shimmer or metallic edge. This artefact sat in the high frequencies. A telephone connection discards nearly everything above roughly 3.4 kHz, which removes the artefact along with the evidence. More on why phone audio is the hard case.
The cues that still weakly work
A few things survive, but read them as mild suspicion rather than as findings. None is strong enough to act on alone, and none is safe to rely on when the alternative is a free phone call.
- The conversation does not adapt. Ask something specific and unexpected — not a security question, a shared memory. A pre-generated clip cannot answer at all, and a live operator running a clone has to improvise around it.
- Unnatural turn-taking. Slight delays before each reply, or a caller who talks over you and never notices. This is as often a bad connection as it is synthesis.
- Emotional tone that does not track the content. Distress that stays at a constant intensity through the whole call, or panic that does not react when you offer help.
- The script does the work. Urgency, secrecy, an unusual payment route, a reason you must not hang up. This is a far better signal than anything acoustic, and it is the one thing that has not changed in decades of fraud.
Notice that the strongest item on that list is not about the voice at all. That is the pattern worth internalising: the content and structure of the request carry more information than the audio does.
Why the procedural answer wins
A cloned voice can imitate a person. It cannot answer a telephone you dialled. That asymmetry is the whole defence, and unlike a listening test it does not degrade as the models improve.
So the rule is: end the call, then dial the number you already had stored — not a number the caller gave you, not a number in the caller ID, not one from a search result you found while still on the line. If the person is genuinely unreachable, contact someone else who would know where they are. If money is involved, confirm it through a channel you initiated yourself.
A household version of this works well: a short passphrase agreed in person, never sent by message and never stored in a phone, that a caller claiming an emergency must be able to produce. It costs nothing and it survives any improvement in synthesis. There is a generator for one here, and a printable card for putting the rule somewhere it will be seen.
For companies the equivalent is a written payment-verification policy, so that no individual employee has to be the one who doubted a senior voice. A template is here.
Where a detector fits
Analysis happens after the fact, on a file. It cannot interrupt a call in progress, and it does not replace the callback. What it can do is help you work out what happened afterwards — whether to warn other people, whether the voice belonged to someone you know, whether this was a clone at all.
Read the result carefully when it comes back. Likely synthetic is the stronger of the two verdicts, because it requires positive evidence: something in the file matched a generation signature. Likely human is weaker, because it can also mean the evidence was destroyed in transit — a forwarded, re-compressed, noise-reduced clip can lose every trace of synthesis and still come back clean. Reading your result goes through this properly.
Questions people ask
Can you tell if a voice is AI just by listening?
Usually not. And the attempt keeps you on the line while you deliberate, which is the condition the call was designed to create. Hang up and dial back on a number you already have.
Do the old tells — robotic pauses, no breathing — still work?
Rarely. They described a generation of systems that has been superseded. On a phone line the codec removes most of what would have remained anyway.
What about asking a trick question?
Better than listening, because it tests knowledge rather than sound. Make it a shared memory rather than a security question — security answers can be researched, and a live operator will simply guess. An agreed passphrase is stronger still.
If a detector says likely human, is the call safe?
No. It means no generation signature was found in that file. A poor recording, a forwarded copy, or a system we have not seen would all produce the same answer. Treat it as one data point, not clearance.
[VERIFY: written by — real name, real role]. Reviewed [VERIFY: date]. Corrections: method@aivoicedetctor.com.
Reviewed