What we measure and how
Accuracy claims in this category are mostly unfalsifiable, because the dataset behind them is never described. This section exists so ours can be argued with.
How our numbers are produced
- Held out, always. Evaluation clips never appear in training. A rate measured on data the model has seen is a memory test, not an accuracy test.
- Balanced across conditions, not just across labels. A test set that is half synthetic and half human but entirely studio-quality tells you nothing about the phone call you actually received.
- Real human speech from real channels. The false positive rate is measured against genuine recordings carrying genuine defects — background noise, clipping, cheap microphones, compression — because that is where false positives come from.
- Re-measured after every retrain. Rates on this site carry a date. When they move, the change is recorded in the changelog, including when a rate gets worse.
The composition of each set, with counts: dataset. Current rates: accuracy.
Open problems we have not solved
Listing these costs us nothing we were entitled to keep, and it is how you tell a research page from a marketing page.
- Re-recorded audio. Playing a generated clip through a speaker and capturing it on a phone gives the file a genuine recording chain, defeating the signal we lean on hardest. We do not have a good answer to this.
- Real-time voice conversion. A live human speaking through a conversion model carries real prosody and real timing. It is a fundamentally harder case than text-to-speech and our rates on it are lower.
- Attribution decay. Naming the generator degrades within weeks of a vendor shipping a new model. The verdict holds up far better than the attribution, and we would rather say that than let people over-read the generator field.
- Languages and accents. Detection is measured mostly on [VERIFY: languages]. Performance outside that is [VERIFY: unmeasured / lower — say which]. Unmeasured is an honest answer; implying uniform performance is not.
- Very short clips. Below [VERIFY] seconds the analysis has too little to work with and confidence should be read accordingly.
Working with us
Researchers. [VERIFY: state whether you provide API access for academic evaluation, on what terms, and who to contact. If the answer is not yet, say so.]
Journalists. We will talk about what detection cannot do as readily as what it can: press@aivoicedetctor.com
Anyone with a clip that beat us. This is the most useful thing you can send: coverage@aivoicedetctor.com. It goes into the next evaluation set, and if it changes a published rate the change appears in the changelog.
Reviewed