When Truthring cannot reliably tell
Detection degrades on short clips, heavy compression, background noise, and audio re-recorded through a speaker — which gives a generated file a genuine recording chain and defeats the strongest signal available. A system newer than the last retrain may still be detected while the generator cannot be named.
This page exists because you should know where the instrument fails before you rely on it — not after. Every condition below measurably degrades the result, and Truthring will usually return unclear rather than guess.
The short answer. Truthring is least reliable on short clips, heavily compressed or re-recorded audio, overlapping speakers, and generation systems released after our last training run. In those conditions treat any verdict — including likely human — as inconclusive.
Conditions that degrade the result
| Condition | Effect | What we do |
|---|---|---|
| Clips under 5 seconds | Too little signal for stable measurement | Confidence capped; usually returns unclear |
| Heavy or repeated compression | Codecs discard the high-frequency detail detection depends on | Confidence reduced in proportion to measured compression depth |
| Re-recorded audio | Playing a clip through a speaker and recording it destroys generation artefacts | Flagged where detectable; result marked low confidence |
| Multiple overlapping speakers | Vocal-tract consistency cannot be attributed to one speaker | Speaker count reported; unclear returned above one |
| Music or heavy background noise | Masks the spectral structure being measured | Signal-to-noise reported; confidence reduced |
| Audio edited after generation | Normalisation, EQ and noise reduction alter or remove signatures | Often undetectable — a known blind spot |
| Generators newer than our last training run | No signature held for the system | May return likely human. See below. |
| Non-English speech | [VERIFY: state your measured performance by language, or say plainly that it is untested] | [VERIFY] |
| Very old or analogue recordings | Degradation resembles compression artefacts | Confidence reduced |
The asymmetry you need to understand
"Likely human" is the weaker verdict. It means we found no signature we recognise. That is not the same as establishing the recording is genuine.
A voice cloned with a system released last week, or a synthetic clip cleaned up after generation, can return likely human. Absence of evidence is not evidence of absence, and we would rather say so on this page than let you discover it in a situation that matters.
"Likely synthetic" is the stronger verdict, because it rests on something found rather than something missing. But it still does not establish who made the clip, when, or why.
What Truthring should not be used for alone
- Legal proceedings. A verdict is a screening signal. Audio submitted as evidence needs forensic examination by a qualified examiner, chain of custody, and provenance analysis.
- Employment decisions. Do not dismiss, discipline or refuse to hire on a detection score.
- Financial transfers. If a voice is asking for money, call the person back on a number you already have. The detector is slower and less certain than that phone call.
- Publication. Screen the audio, then verify source, timestamp and corroborating evidence before you publish.
- Accusing someone. A false positive on a genuine recording is a real and measurable outcome. So is a false negative.
Known blind spots
Stated plainly, because a detector that claims none is not being honest with you.
- Post-generation processing. Deliberate cleanup after synthesis can remove signatures we rely on. We do not currently detect this reliably.
- Novel architectures. A genuinely new synthesis approach may produce no signature we hold until we have trained against it.
- Adversarial audio. Someone who knows what a detector measures can degrade a clip specifically to defeat it.
- [VERIFY: any blind spot specific to your engine. Add rather than remove — this list is the page's whole value.]
How to read a low-confidence result
Confidence below [VERIFY: your threshold] means the recording did not carry enough usable signal. That is a statement about the clip, not about the speaker.
If you can obtain a longer or cleaner version of the same audio — the original file rather than a forwarded copy, or a section without background noise — run that instead. Ten seconds of clean speech will tell you more than three minutes of a compressed forward.
[VERIFY: Written by — real name, real role.] Reviewed [VERIFY: date]. Engine v2.4.
If you have found a failure mode not listed here, tell us and we will add it: method@aivoicedetctor.com.
Questions
What are the limitations of AI voice detection?
Short clips, heavy compression, background noise and re-recording through a speaker all degrade a result. Detection on phone audio is materially weaker than on clean files because a phone network discards most audio above roughly 3.4 kHz, which is where much of the evidence sits.
Is a likely human result the same as a clean bill of health?
No, and this is the asymmetry that matters most. Likely synthetic is the strong verdict because it requires positive evidence. Likely human is weak, because it can also mean the evidence was destroyed by compression in transit. On a compressed phone call it is close to no information at all.
Can a Truthring result be used as proof?
No. It is a probability produced by a stated method with a stated error rate, usable as one input alongside others and unusable on its own. It must never be the sole basis for accusing, disciplining, dismissing or prosecuting anyone.
Reviewed