AI voices on a phone call
Yes, synthetic speech can be detected in a recorded phone call — less reliably than in a clean file, and the gap is large enough that it should change how you read the result. Truthring reports [VERIFY]% on recorded calls against [VERIFY]% on studio-quality audio. This page explains where the difference goes.
Most voice fraud arrives by phone. That is not a coincidence — it is the channel where detection is weakest and urgency is easiest to manufacture. Any vendor quoting one accuracy figure without saying which condition it was measured in is quoting the studio number.
What a phone network removes
A telephone call is not a recording of a voice. It is a reconstruction, built to be intelligible over a narrow channel at low cost. To achieve that, the network throws away almost everything that is not needed to understand words.
- Everything above roughly 3.4 kHz. Traditional telephony carries a band of about 300 Hz to 3.4 kHz. Human hearing goes to around 20 kHz. The upper region — where breath, sibilance and the fine texture of a real recording live — simply is not transmitted.
- The noise floor. Codecs are designed to suppress background, and modern ones aggressively remove what they judge to be non-speech. The traces of a real room and a real microphone are exactly what they are built to discard.
- Fine timing detail. Packet-based calls reassemble audio from fragments, conceal lost packets by interpolating, and re-time what arrives late. The result is intelligible and no longer a faithful record of what was said.
Put plainly: the phone network is a filter that strips out most of what distinguishes captured audio from manufactured audio, and it applies that filter equally to a real caller and a cloned one.
What survives, and what we measure
Not everything is lost. Detection on phone audio leans on the parts of speech that live inside the transmitted band and that a generator still has to get right.
Prosody across a whole turn
How emphasis, pace and pitch develop over several sentences rather than within one. Long-range structure survives compression far better than fine spectral detail, which is why a longer clip helps disproportionately here.
Behaviour at interruptions
Real conversation overlaps, stumbles, restarts and abandons sentences. Generated speech recovers from interruption in a way that is measurably tidier, and tidiness is not removed by a codec.
Breath and pause distribution
Where a speaker takes a breath, and how the length of pauses is distributed across a turn. Attenuated by the codec, not eliminated.
Consistency with the channel
Whether the audio is consistent with having actually passed through a call, or with having been generated and then filtered to sound as if it had. These leave different marks.
How to read a verdict on a call recording
The asymmetry matters more here than anywhere else on this site.
- “Likely synthetic” on phone audio is meaningful. The evidence had to survive the codec to be found, so finding it is a real signal.
- “Likely human” on phone audio is weak. It can mean the speech was human. It can equally mean the evidence of synthesis was destroyed in transit. Do not treat it as clearance.
- “Unclear” is the honest outcome for many calls, particularly short ones. We would rather return it than manufacture confidence the audio does not support.
Detection rates by condition, including the rate at which real human speech is wrongly flagged: accuracy.
Getting the best recording you can
- Keep the original file. Every forward through a messaging app re-compresses it. Send us the file the recorder produced, not a copy of a copy.
- Submit the longest continuous stretch of speech, not the most incriminating sentence. Length is the single biggest lever you control on phone audio.
- Do not clean it up. Noise reduction and normalisation remove evidence along with the noise. Send it raw.
- Do not re-record it by playing it on a speaker and capturing it on another phone. That destroys the strongest remaining signal.
- Note the channel. A carrier call, a WhatsApp call and a Zoom recording compress differently, and telling us which one it was improves how we read the result.
What to do during the call, before any of this matters
No detector helps in the moment, ours included. The countermeasure that works is procedural and costs nothing.
Hang up and call back on a number you already have stored. Not a number the caller gives you. A voice clone cannot answer a phone you dialled, and essentially every voice fraud that succeeds does so because the target stayed inside the channel the attacker chose. Urgency is the tell: a real emergency survives a two-minute callback, and a manufactured one does not.
Questions
Can AI voices be detected on a phone call?
Yes, but less reliably than on a clean file. A phone network discards most audio above roughly 3.4 kHz, which is where much of the evidence of synthesis sits. Truthring detects synthetic speech in recorded calls at a rate not yet measured, against a rate not yet measured on studio-quality audio.
Can you analyse a call while it is happening?
No. Truthring works on a recorded file after the fact. If a call feels wrong while it is in progress, hang up and dial the person back on a number you already have.
Is it legal for me to record the call?
It depends on where you and the other party are. Some places require only one party to consent; others require everyone on the line. Check the rule that applies to you before recording — an unlawfully made recording is usually worthless in the proceeding you wanted it for, and can create a problem of its own.
Does a voicemail work better than a live call?
Usually yes. Voicemail is stored after a single compression pass and is often a longer uninterrupted stretch of speech, both of which help. A voicemail is the best phone-derived audio you are likely to have.
Rates on this page measured on [VERIFY: n] call recordings across [VERIFY: n] carriers and app channels, re-measured [VERIFY: date].
Reviewed