How accurate is Truthring?
Truthring will publish both error directions — the detection rate and the rate at which genuine human speech is wrongly flagged — broken out by recording condition, with the dataset described and the date attached. No measured figures exist yet. The evaluation framework is published in advance so the method can be challenged before any number exists.
Detection rates by recording condition, with the sample size stated for each. Measured on the kind of audio people actually submit — compressed, noisy, forwarded — not on studio recordings.
The short answer. [VERIFY: one sentence, e.g. "Truthring correctly identifies synthetic speech in X% of telephone-quality clips and misclassifies genuine speech as synthetic in Y% of cases."] Accuracy falls as recording quality falls, and the table below shows by how much.
Results by condition
| Condition | Samples | Detection rate | False positive rate |
|---|---|---|---|
| Clean synthetic speech Studio quality, uncompressed | [VERIFY] | [VERIFY]% | — |
| Telephone-quality synthetic 8 kHz, single codec pass | [VERIFY] | [VERIFY]% | — |
| Doubly compressed synthetic Forwarded voice note | [VERIFY] | [VERIFY]% | — |
| Synthetic with background noise | [VERIFY] | [VERIFY]% | — |
| Genuine human speech, clean | [VERIFY] | — | [VERIFY]% |
| Genuine human speech, telephone | [VERIFY] | — | [VERIFY]% |
Measured [VERIFY: month/year] on engine v2.4. Test set described below. Results vary with recording quality, codec, speaker, language and generation system.
The number most detectors don't publish
Our false positive rate on genuine human speech is [VERIFY]%.
That is the rate at which real recordings get flagged as synthetic. It matters more than the headline detection rate, because a false positive is what wrongly accuses someone — and in this category that is the expensive kind of error.
Any detector quoting a single "99% accurate" figure without separating detection rate from false positive rate is not telling you enough to use the number.
What we tested on
- Synthetic samples: [VERIFY: how many, from which generation systems, generated when]
- Human samples: [VERIFY: how many, from which sources, and whether speakers gave consent]
- Degradation: [VERIFY: which codecs and bitrates you passed audio through, and how noise was added]
- Held out: [VERIFY: confirm the test set was not used in training — this is the sentence a reviewer looks for]
The test set is described in full at research/dataset. [VERIFY: state whether it is available on request, and to whom.]
Why our clean-audio numbers are lower than some competitors'
They are, and it is deliberate. Truthring's model is weighted toward telephone-quality audio, because that is what arrives. A model tuned to score well on studio samples performs worse on a forwarded WhatsApp voice note — which is the file people actually bring to a detector.
If your use case is clean studio audio, a detector optimised for that will beat us. We would rather say so than quote a number measured on conditions you will never encounter.
How to read these numbers
A detection rate is a population statistic. It tells you how the model performs across a test set; it does not tell you how confident to be about one specific clip. That is what the confidence figure on an individual verdict is for.
A 90% detection rate means one clip in ten is missed. If the consequence of missing one matters, the detector is a screening step and not the decision.
[VERIFY: Written by — real name, real role.] Measured [VERIFY: date]. Engine v2.4.
Challenges to this methodology are welcome: method@aivoicedetctor.com.
Questions
How accurate is Truthring?
No measured figure exists yet and none will be published until it does. The evaluation framework is on this page in advance of results, including the false positive rate, which most detectors omit. Any figure on this site that is not yet measured is marked pending.
Why publish the false positive rate at all?
A detection rate on its own can be raised simply by flagging more aggressively; what that hides is the rate at which genuine human speech gets flagged as fake. The two move in opposite directions. For anyone being accused on the strength of a result, the false positive rate is the number that matters.
Why are clean-audio figures lower than some competitors quote?
Because a rate measured on studio-quality audio does not describe the voicemail or call recording most people actually submit. Rates here are broken out by condition rather than reported as a single headline number.
Reviewed