Truthring
Detection · the main tool

What an AI voice detector is, and what this one does

In short

An AI voice detector is software that examines a recording and estimates whether the speech in it was generated by a machine or captured from a person. It returns a probability rather than a ruling. Truthring publishes that probability together with the rate at which genuine human speech is wrongly flagged.

Most tools in this category answer a single question and answer it loudly. The harder and more useful job is telling you how much the answer is worth on the particular file you have, which depends on how long it is, what it has been through, and whether the generator behind it is one anyone has seen before.


What Truthring does

It takes one audio file and looks for the difference between speech that was captured and speech that was rendered. Captured speech carries the fingerprints of a physical chain: a room with walls, a body producing air pressure, a microphone with its own noise floor and its own imperfections. Rendered speech is written straight to a buffer. A generator has no reason to reproduce that chain faithfully, because sounding right to a listener was the entire objective and reproducing a microphone was not.

Alongside the verdict comes a confidence figure, the generator’s name where the clip matches a signature we hold, and a reference code. The reference is the least glamorous field and the most important one: it is what lets a result be pulled up again months later, after the model behind it has been retrained, and checked against the method that produced it.

The verdicts are also asymmetric, and reading them as a matched pair will lead you wrong. A result of likely synthetic requires evidence to appear at all. A result of likely human can mean the speech was human, or it can mean whatever would have given a generator away did not survive the trip through a phone network. A clean studio file that reads human is informative. A twice-forwarded voice note that reads human is close to silence.

What Truthring deliberately does not do

It does not listen to live calls. Streaming audio through a detector while two people are talking sounds like the obvious product and is not one: call audio is aggressively compressed in transit, and a verdict delivered mid-sentence is a verdict nobody can check afterwards. We analyse a recorded file, after the fact, with the file still available for a second opinion.

It does not identify speakers. A result says something about how a recording was made, never about who made it. Truthring will not tell you that a voice belongs to a named person, and a synthetic verdict is not an accusation against whoever sent you the file.

It does not settle whether a recording is honest. Audio can be wholly authentic and still leave a listener with a false impression, which is a problem of editing rather than of synthesis and needs a different kind of examination. That distinction gets its own page: deepfake audio detection.

And it does not sell certainty. There is no threshold at which a probability becomes a fact, no badge, and no wording anywhere on this site that would let a result be waved at someone as proof.


What happens when you upload a clip

The flow at /check/ is deliberately short. Nothing about it requires an account.

01You drop a file, or choose one from diskseconds
02Format and length are checked before anything is uploadedlocal
03Speech is isolated and the recording chain is examinedanalysis
04The clip is compared against generator signaturesanalysis
05A verdict, a confidence figure and a reference code come backresult

The prototype interface accepts MP3, WAV, M4A, OGG and WEBM. For comparison, TextSight — the AI text and voice detection suite run by the same company — ships a working voice detector that accepts MP3, WAV, M4A, OGG and FLAC up to 10 MB. Truthring’s own ceiling is still being set: [VERIFY: Truthring file limits at launch]


Four things that change the answer

Two people can submit clips of the same generated voice and get results of very different strength. Almost all of the variation comes from four properties of the file rather than from the model.

Length

Short clips carry less evidence

Signature analysis and prosody analysis both need material to work with. A four-second clip leans almost entirely on the recording-chain pass, and the confidence figure will say so. Send the longest continuous stretch of speech you have, not the most incriminating sentence in it.

Compression

Codecs remove exactly what we read

Phone networks and messaging apps discard fine detail on purpose, because the human ear does not miss it. That fine detail is where the difference between captured and rendered audio lives. Every re-encode — every forward through another app — removes more of it. Check the original file if you can still reach it.

Re-recording

Playing a clip aloud gives it a real recording chain

Generated audio played through a speaker and captured on a phone genuinely passed through a room and a microphone. The first pass is now looking at truthful evidence of a physical capture, and it will report what it sees. This is the cheapest way to defeat the analysis and it needs no technical skill at all.

Noise

Background sound masks the signal in both directions

Traffic, a crowd, music under speech. Noise can bury the traces that would expose a generator, and it can also make a genuine recording look manufactured by swamping the natural detail we expect to find. It pushes results towards unclear, which is the correct place for them to go.

The full catalogue, including the conditions we have not yet measured, is on limitations.


What a result looks like

Specimen only. The figures below are placeholders illustrating the layout, not the output of any analysis.

Specimen · synthetic

Likely synthetic

[VERIFY]
Confidence

A short plain-language reason sits here, naming what was found rather than restating the verdict. The rows below carry everything a second reader would need.

Attributed generator
Specimen — may read “unknown generator”
False positive rate, this channel
[VERIFY: measured rate]
Source quality
Phone-compressed · 0:14
Model version
[VERIFY: version at launch]
Reference
TR-8F29A1

How to read each field, line by line: reading your result.


Which generators we can name

Naming a system and detecting one are separate achievements. A clip can be correctly called synthetic while the generator stays unknown, which happens whenever the system that made it is newer than our last training run. Attribution is the bonus. The verdict is the product.

Per-generator detection and attribution figures are published once the first benchmark run completes: [VERIFY: per-generator measured rates pending first benchmark]

Voice cloning

Systems that rebuild a named individual’s voice from a short recording of them. These are the ones that turn up in fraud.

SystemWhat it producesMeasured rates
ElevenLabsCloned speech from a short sampleForthcoming
Resemble AICloned speech, real-time and batchForthcoming
PlayHTCloned and stock voicesForthcoming
Descript OverdubCloned speech inside an editorForthcoming
LovoCloned and stock voicesForthcoming
Replica StudiosCloned performance voicesForthcoming
Camb.aiCloned speech across languagesForthcoming

Text to speech

Systems that read text aloud in a stock or designed voice rather than copying a named person.

SystemWhat it producesMeasured rates
OpenAI TTSStock voicesForthcoming
Azure Neural TTSStock voicesForthcoming
Google WaveNetStock voicesForthcoming
Amazon PollyStock voicesForthcoming
MurfStock voicesForthcoming
SpeechifyStock voicesForthcoming
WellSaid LabsStock voicesForthcoming
iSpeechStock voicesForthcoming
Cartesia SonicLow-latency stock voicesForthcoming
Deepgram AuraLow-latency stock voicesForthcoming
RimeConversational stock voicesForthcoming
Hume AIExpressive stock voicesForthcoming
Fish AudioStock and cloned voicesForthcoming

Open weight and research

Models that run on a laptop. No vendor account, no rate limit, no log of who generated what.

SystemWhat it producesMeasured rates
Coqui XTTSOpen-weight cloningForthcoming
BarkOpen-weight generative speechForthcoming
Tortoise TTSOpen-weight cloningForthcoming
VALL-E / VALL-E XResearch cloning from seconds of audioForthcoming

Music and song

Sung vocals behave differently from speech, and results here are not comparable with the tables above.

SystemWhat it producesMeasured rates
SunoGenerated songs with vocalsForthcoming
UdioGenerated songs with vocalsForthcoming

Open-weight systems are a different problem. A model anyone can download has no vendor, no consent checkbox and no release calendar, which changes both who uses it and how quickly a signature stops matching. Treated as a category here: open-weight voice cloning.

When the generator is one we do not hold

The verdict reads unknown generator and stops there. The synthetic-or-human judgement does not depend on recognising a vendor, so it can still be correct and still be useful. What you lose is the name, and the name is the part that decays fastest, because vendors ship new models on their own schedule and every release moves the signature a little.

If you have a clip from a system that produced a result you believe is wrong, send it: coverage@aivoicedetctor.com. Clips that break us are more valuable than clips that confirm us.


Questions

What is an AI voice detector?

Software that examines a recording and estimates whether the speech in it was manufactured by a model or captured from a person through a microphone. It returns a probability with a stated error rate, not a yes or a no, and it works on the file rather than on the claim attached to it.

Can an AI voice detector be wrong?

Routinely, in both directions. It can miss a well-made clone, and it can flag a real person recorded badly. That is why Truthring publishes the rate at which genuine speech is wrongly flagged next to the rate at which synthetic speech is caught. A tool that shows you only the flattering number is hiding half the result.

Why is “likely human” the weaker verdict?

Because there are two ways to reach it. The clip really was spoken by a person, or the evidence that would have exposed a generator was destroyed before the file reached us — by a phone codec, by re-encoding, by noise. “Likely synthetic” needs positive evidence to appear, so it carries more weight.

What audio formats can I submit?

The prototype at /check/ accepts MP3, WAV, M4A, OGG and WEBM. TextSight, the sibling product from the same company, runs a live voice detector that takes MP3, WAV, M4A, OGG and FLAC up to 10 MB. Truthring’s own ceiling at launch is not fixed yet.

Will the result name which system generated the clip?

Only when the audio matches a signature we hold. Otherwise the verdict reads unknown generator, and stays there. Guessing at the nearest match would produce a confident wrong name, which is worse in a dispute than an honest blank.

Reviewed