What an AI voice detector is, and what this one does
An AI voice detector is software that examines a recording and estimates whether the speech in it was generated by a machine or captured from a person. It returns a probability rather than a ruling. Truthring publishes that probability together with the rate at which genuine human speech is wrongly flagged.
Most tools in this category answer a single question and answer it loudly. The harder and more useful job is telling you how much the answer is worth on the particular file you have, which depends on how long it is, what it has been through, and whether the generator behind it is one anyone has seen before.
What Truthring does
It takes one audio file and looks for the difference between speech that was captured and speech that was rendered. Captured speech carries the fingerprints of a physical chain: a room with walls, a body producing air pressure, a microphone with its own noise floor and its own imperfections. Rendered speech is written straight to a buffer. A generator has no reason to reproduce that chain faithfully, because sounding right to a listener was the entire objective and reproducing a microphone was not.
Alongside the verdict comes a confidence figure, the generator’s name where the clip matches a signature we hold, and a reference code. The reference is the least glamorous field and the most important one: it is what lets a result be pulled up again months later, after the model behind it has been retrained, and checked against the method that produced it.
The verdicts are also asymmetric, and reading them as a matched pair will lead you wrong. A result of likely synthetic requires evidence to appear at all. A result of likely human can mean the speech was human, or it can mean whatever would have given a generator away did not survive the trip through a phone network. A clean studio file that reads human is informative. A twice-forwarded voice note that reads human is close to silence.
What Truthring deliberately does not do
It does not listen to live calls. Streaming audio through a detector while two people are talking sounds like the obvious product and is not one: call audio is aggressively compressed in transit, and a verdict delivered mid-sentence is a verdict nobody can check afterwards. We analyse a recorded file, after the fact, with the file still available for a second opinion.
It does not identify speakers. A result says something about how a recording was made, never about who made it. Truthring will not tell you that a voice belongs to a named person, and a synthetic verdict is not an accusation against whoever sent you the file.
It does not settle whether a recording is honest. Audio can be wholly authentic and still leave a listener with a false impression, which is a problem of editing rather than of synthesis and needs a different kind of examination. That distinction gets its own page: deepfake audio detection.
And it does not sell certainty. There is no threshold at which a probability becomes a fact, no badge, and no wording anywhere on this site that would let a result be waved at someone as proof.
What happens when you upload a clip
The flow at /check/ is deliberately short. Nothing about it requires an account.
The prototype interface accepts MP3, WAV, M4A, OGG and WEBM. For comparison, TextSight — the AI text and voice detection suite run by the same company — ships a working voice detector that accepts MP3, WAV, M4A, OGG and FLAC up to 10 MB. Truthring’s own ceiling is still being set: [VERIFY: Truthring file limits at launch]
Four things that change the answer
Two people can submit clips of the same generated voice and get results of very different strength. Almost all of the variation comes from four properties of the file rather than from the model.
Short clips carry less evidence
Signature analysis and prosody analysis both need material to work with. A four-second clip leans almost entirely on the recording-chain pass, and the confidence figure will say so. Send the longest continuous stretch of speech you have, not the most incriminating sentence in it.
Codecs remove exactly what we read
Phone networks and messaging apps discard fine detail on purpose, because the human ear does not miss it. That fine detail is where the difference between captured and rendered audio lives. Every re-encode — every forward through another app — removes more of it. Check the original file if you can still reach it.
Playing a clip aloud gives it a real recording chain
Generated audio played through a speaker and captured on a phone genuinely passed through a room and a microphone. The first pass is now looking at truthful evidence of a physical capture, and it will report what it sees. This is the cheapest way to defeat the analysis and it needs no technical skill at all.
Background sound masks the signal in both directions
Traffic, a crowd, music under speech. Noise can bury the traces that would expose a generator, and it can also make a genuine recording look manufactured by swamping the natural detail we expect to find. It pushes results towards unclear, which is the correct place for them to go.
The full catalogue, including the conditions we have not yet measured, is on limitations.
What a result looks like
Specimen only. The figures below are placeholders illustrating the layout, not the output of any analysis.
Likely synthetic
A short plain-language reason sits here, naming what was found rather than restating the verdict. The rows below carry everything a second reader would need.
- Attributed generator
- Specimen — may read “unknown generator”
- False positive rate, this channel
- [VERIFY: measured rate]
- Source quality
- Phone-compressed · 0:14
- Model version
- [VERIFY: version at launch]
- Reference
- TR-8F29A1
How to read each field, line by line: reading your result.
Which generators we can name
Naming a system and detecting one are separate achievements. A clip can be correctly called synthetic while the generator stays unknown, which happens whenever the system that made it is newer than our last training run. Attribution is the bonus. The verdict is the product.
Per-generator detection and attribution figures are published once the first benchmark run completes: [VERIFY: per-generator measured rates pending first benchmark]
Voice cloning
Systems that rebuild a named individual’s voice from a short recording of them. These are the ones that turn up in fraud.
| System | What it produces | Measured rates |
|---|---|---|
| ElevenLabs | Cloned speech from a short sample | Forthcoming |
| Resemble AI | Cloned speech, real-time and batch | Forthcoming |
| PlayHT | Cloned and stock voices | Forthcoming |
| Descript Overdub | Cloned speech inside an editor | Forthcoming |
| Lovo | Cloned and stock voices | Forthcoming |
| Replica Studios | Cloned performance voices | Forthcoming |
| Camb.ai | Cloned speech across languages | Forthcoming |
Text to speech
Systems that read text aloud in a stock or designed voice rather than copying a named person.
| System | What it produces | Measured rates |
|---|---|---|
| OpenAI TTS | Stock voices | Forthcoming |
| Azure Neural TTS | Stock voices | Forthcoming |
| Google WaveNet | Stock voices | Forthcoming |
| Amazon Polly | Stock voices | Forthcoming |
| Murf | Stock voices | Forthcoming |
| Speechify | Stock voices | Forthcoming |
| WellSaid Labs | Stock voices | Forthcoming |
| iSpeech | Stock voices | Forthcoming |
| Cartesia Sonic | Low-latency stock voices | Forthcoming |
| Deepgram Aura | Low-latency stock voices | Forthcoming |
| Rime | Conversational stock voices | Forthcoming |
| Hume AI | Expressive stock voices | Forthcoming |
| Fish Audio | Stock and cloned voices | Forthcoming |
Open weight and research
Models that run on a laptop. No vendor account, no rate limit, no log of who generated what.
| System | What it produces | Measured rates |
|---|---|---|
| Coqui XTTS | Open-weight cloning | Forthcoming |
| Bark | Open-weight generative speech | Forthcoming |
| Tortoise TTS | Open-weight cloning | Forthcoming |
| VALL-E / VALL-E X | Research cloning from seconds of audio | Forthcoming |
Open-weight systems are a different problem. A model anyone can download has no vendor, no consent checkbox and no release calendar, which changes both who uses it and how quickly a signature stops matching. Treated as a category here: open-weight voice cloning.
When the generator is one we do not hold
The verdict reads unknown generator and stops there. The synthetic-or-human judgement does not depend on recognising a vendor, so it can still be correct and still be useful. What you lose is the name, and the name is the part that decays fastest, because vendors ship new models on their own schedule and every release moves the signature a little.
If you have a clip from a system that produced a result you believe is wrong, send it: coverage@aivoicedetctor.com. Clips that break us are more valuable than clips that confirm us.
Questions
What is an AI voice detector?
Software that examines a recording and estimates whether the speech in it was manufactured by a model or captured from a person through a microphone. It returns a probability with a stated error rate, not a yes or a no, and it works on the file rather than on the claim attached to it.
Can an AI voice detector be wrong?
Routinely, in both directions. It can miss a well-made clone, and it can flag a real person recorded badly. That is why Truthring publishes the rate at which genuine speech is wrongly flagged next to the rate at which synthetic speech is caught. A tool that shows you only the flattering number is hiding half the result.
Why is “likely human” the weaker verdict?
Because there are two ways to reach it. The clip really was spoken by a person, or the evidence that would have exposed a generator was destroyed before the file reached us — by a phone codec, by re-encoding, by noise. “Likely synthetic” needs positive evidence to appear, so it carries more weight.
What audio formats can I submit?
The prototype at /check/ accepts MP3, WAV, M4A, OGG and WEBM. TextSight, the sibling product from the same company, runs a live voice detector that takes MP3, WAV, M4A, OGG and FLAC up to 10 MB. Truthring’s own ceiling at launch is not fixed yet.
Will the result name which system generated the clip?
Only when the audio matches a signature we hold. Otherwise the verdict reads unknown generator, and stays there. Guessing at the nearest match would produce a confident wrong name, which is worse in a dispute than an honest blank.
How it works
How well it works
Related detection
Reviewed