False positive rate
The number this site is built around. Every detector has one; publishing it is a choice, and a headline accuracy figure quoted without it cannot be used for anything.
A false positive rate is the proportion of genuine human recordings that a detector wrongly flags as synthetic. It is measured on real speech only, and it moves independently of the detection rate. Quoted on its own, a detection rate says nothing about how often a system will accuse someone who did nothing.
Two errors, measured on two different piles
A detector can be wrong in two directions, and the mistake people make is treating them as one quantity.
The detection rate is measured on synthetic samples: of the generated clips we put in, how many were caught. The false positive rate is measured on human samples: of the genuine clips we put in, how many were flagged anyway. Nothing about the first tells you the second, because they are computed on separate sets of audio.
They also trade against each other. Every detector has a threshold, and moving it improves one number while worsening the other. A system can be tuned to catch essentially everything synthetic, at the cost of flagging a great deal of real speech, and the resulting detection rate looks magnificent in a headline. A single blended “accuracy” figure hides all of this, because its value depends on how many synthetic and genuine samples the vendor chose to put in the test set — a ratio nobody outside the vendor can see.
Why the rate feels smaller than it behaves
The arithmetic below is illustrative. It uses invented figures to show a structural effect, and is not a claim about any product, ours included.
Imagine a review queue where one clip in a hundred is genuinely synthetic. Put a thousand clips through a detector with a 90% detection rate and a 5% false positive rate. Ten clips are synthetic, and nine are caught. Nine hundred and ninety are genuine, and about fifty are flagged anyway. The reviewer sees fifty-nine flags, of which nine are right.
Most of what lands in front of that reviewer is innocent, and no improvement to the detection rate fixes it, because the flood is coming from the other number. This is why a low-sounding false positive rate has to be read against how rare the thing being detected actually is in the population you are running it on. It is also why an escalation process that treats a flag as a finding will spend most of its time on people who did nothing.
A concrete case
A university receives an audio assignment submitted by a student who recorded it on a cheap headset in a hard-surfaced room, then compressed it to fit an upload limit. It is flagged as synthetic. She did not use a generator; the recording is simply degraded in ways that resemble what degradation does to a signal.
She now has to prove a negative to someone holding a printout with a percentage on it. The cost of that error is not shared equally with the institution that made it, which is the reason we treat the false positive rate as the more important of the two figures rather than the one that spoils the marketing.
Commonly confused with: the chance a flag is wrong
These are different quantities and the difference is the whole of the previous section. The false positive rate answers: given that this recording is genuine, how often is it flagged? The question a reviewer actually has is the reverse: given that this recording was flagged, how likely is it to be genuine? The second depends on the first and on how common synthetic audio is in the material being reviewed.
It is also distinct from the false negative rate, which is the share of synthetic clips that slip through, and from a system’s verification error rates, which are about identity and are measured on an entirely different task. Borrowing a number from one of these to describe another is common and always misleading.
What Truthring publishes
Both directions, broken down by recording condition, with the sample size stated for each row, on the accuracy page. Clean audio and telephone audio are reported separately because the difference between them is large and hiding it flatters the result. Our clean-audio figures are lower than some competitors quote, for the reason set out on the phone audio page: the model is weighted toward the compressed, forwarded, noisy files people actually submit.
[VERIFY: measured detection and false positive rates by condition — these remain pending until the evaluation run on the released engine is complete]
The same reasoning shapes the verdicts themselves. Likely synthetic requires positive evidence and is the stronger claim. Likely human is the weaker one, because it is also what a clip returns when the evidence was destroyed before it arrived. And unclear exists so that a poor recording does not have to be forced into one of the other two, which is where a large share of false positives would otherwise be manufactured. The evaluation dataset describes what the figures are measured on.
Questions this term raises
Why does the false positive rate matter more than the detection rate?
Because of what each error costs. A missed synthetic clip means a threat was not caught. A false positive means a real person was wrongly accused, usually with no easy way to disprove it. In this category the second is the expensive error.
What does a single accuracy percentage tell me?
Less than it appears to. It blends both error types at whatever ratio of synthetic to genuine samples the test set happened to contain, and that ratio is a choice made by whoever ran the test. Ask for the two rates separately, with sample sizes.
If a detector flags my recording, does that mean I did something wrong?
No. It means the analysis found something consistent with generation. Poor microphones, heavy compression, noise reduction and reverberant rooms all push genuine audio in that direction. A flag is a reason to look further, not a finding about a person.
Are Truthring’s numbers published yet?
Not in final form. The product is pre-launch, and quantitative results are forthcoming rather than available. What is already fixed is the commitment to publish both directions by recording condition, with sample sizes, rather than a single headline figure.
Reading a result
Reviewed