What is in the test set
Every accuracy figure on this site was measured on the clips described below. Publishing the composition is what makes the figures arguable — and a figure nobody can argue with is not a measurement.
Synthetic clips
| Source system | Clips | Voices | Generated |
|---|---|---|---|
| [VERIFY: generator] | [VERIFY] | [VERIFY] | [VERIFY: date] |
| [VERIFY: one row per generator in /voices/. Generation date matters — a set built a year ago measures a model nobody uses now.] | |||
Human clips
The set behind the false positive rate. It is deliberately awkward: if it were all clean studio speech the false positive rate would look excellent and mean nothing.
| Source | Clips | Speakers | Consent basis |
|---|---|---|---|
| [VERIFY: e.g. public-domain speech corpora, name each] | [VERIFY] | [VERIFY] | [VERIFY] |
| [VERIFY: recordings collected with consent] | [VERIFY] | [VERIFY] | [VERIFY] |
| [VERIFY: state the consent basis for every human recording. A detector that trained on people's voices without permission cannot credibly lecture anyone about voice cloning.] | |||
Channel conditions
| Condition | How produced | Share of set |
|---|---|---|
| Studio | Original file, no re-encoding | [VERIFY]% |
| Voice note | [VERIFY: which apps, which codecs] | [VERIFY]% |
| Phone call | [VERIFY: carriers and VoIP paths used] | [VERIFY]% |
| Noisy environment | [VERIFY] | [VERIFY]% |
| Re-recorded through a speaker | [VERIFY] | [VERIFY]% |
Languages
[VERIFY: list languages with clip counts. Then state plainly what performance outside them is — measured and lower, or unmeasured. Unmeasured is an acceptable answer; silence is not.]
What we will share
[VERIFY: whether the human set, the synthetic set, or the generation scripts can be shared with researchers, and on what terms. If licensing prevents sharing a corpus, name the corpus so someone else can obtain it themselves.]
Reviewed