Truthring
For researchers · method, versions and datasets

You do not need our verdict. You need enough of our method to attack it.

A detector is only useful to a researcher if its claims can be reproduced, contested and cited three years later. That needs four things: a written method, a version string on the model behind each result, a described evaluation set, and error rates published in both directions.


The shortage is not architectures

Synthetic speech detection has plenty of models. What it lacks are claims that survive contact with somebody else’s data. A detector announces a percentage, does not name the corpus, does not state the decision threshold, does not say whether the clips were held out, and never reports what happened to genuine speech. That number is unfalsifiable by construction. The useful question to put to any vendor, us included, is not how accurate but on what, measured how, and what did the human speech do.


What we publish, and in what order

The false positive rate appears wherever the detection rate appears. Genuine speech wrongly flagged is the error that harms people, and any detector can be made to look impressive on synthetic clips by flagging more of everything. One number without the other conceals the trade.

The two verdict states are asymmetric, which matters more for evaluation design than for users. Likely synthetic is a positive finding: something matched a generation artefact. Likely human is a null, and here nulls are contaminated by channel effects — narrowband coding, bitrate reduction, a clip re-encoded three times through a messaging app. A test set without deliberately degraded genuine speech reports a false positive rate that means nothing outside a studio.

Every rate carries an engine version and a date. After a retrain the rates are re-measured and the movement goes into the changelog, including when one gets worse. A result cited without its version is not reproducible, so the version travels with the report. [VERIFY: publish the first measured rates with dataset composition, engine version and measurement date before citing any figure]


The dataset composition is the actual paper

The dataset page lists what the rates were measured on: which generator produced each block of synthetic clips and when, where the genuine speech came from, the consent basis for every human recording, and the channel conditions applied, from studio down to audio re-recorded through a loudspeaker.

Two entries do most of the work. The generation date bounds the claim: a set built against an earlier generation of systems measures models fewer people now use, and every detector flatters itself on one. The consent basis is an ethics question better answered in public than under questioning, since a detector trained on voices harvested without permission is poorly placed to lecture anyone about cloning. Where licensing blocks redistribution the corpus is named so you can obtain it directly — weaker openness than shipping files, and better labelled than blurred.


Nobody independent has audited any of this

There has been no third-party audit of the model, the evaluation pipeline or the published rates, no external replication of the benchmark, and no certification of any kind. Truthring is a product of Lacewing Technologies, a small independent software company in Navi Mumbai, and the evaluation is self-reported by the people who built the thing being evaluated.

The benchmark carries the defect in sharper form: we designed it, we are in it, and we chose the conditions. The only mitigation is publishing the method fully enough that someone with no incentive to flatter us can run it and get a different answer. If your group does, publish it.


What access you can request

There is no formal programme yet. What a research group needs — batch volume that makes an evaluation feasible, scored output rather than verdict labels, a pinned engine version for the length of a study, and permission to publish without review — must be defined before it can be promised. [VERIFY: researcher access programme not yet defined — set terms for academic API volume, score-level output, version pinning and publication rights, then state them here]

Two things already hold. Adversarial examples are welcome: a clip the engine got wrong is worth more than one it got right, and those enter the next evaluation set with the failure recorded rather than quietly patched. And where a published result is ambiguous we will tell you the engine version, the conditions and what the confidence meant.


What this cannot do for your research

It is not open weights. You can reproduce the evaluation procedure; you cannot reproduce the model from what is published. [VERIFY: state whether weights, feature extractors or scoring code will be released, and under what licence] Treat it as a black-box system under study, not a baseline you can retrain.

It cannot serve as ground truth for labelling a corpus. Using one detector to label clips you then evaluate detectors on imports that detector’s error distribution into your labels. If our engine misses a generator family, your ground truth inherits the blind spot and your paper measures agreement with us rather than correctness.

It cannot attribute a clip to a person, a device or a place. Attribution means a generator family whose signature we already hold, and it decays within weeks of a vendor shipping a new model.

It cannot substitute for human adjudication. Where a finding touches a real person or event, the determination belongs to a named human who has weighed provenance, corroboration and context and can be questioned on it. A score is one line of evidence, rarely the strongest.


Citing a result so it stays citable

Five fields make an output checkable by a reviewer: engine version, analysis date, the verdict state including unclear rather than collapsed into a binary, the confidence as reported, and the reference code resolving to the report. Flattening three states into two is the commonest way detector comparisons become incomparable, and it favours whichever tool abstains most. State the audio condition too, since a studio rate and an 8 kHz telephone rate are claims about different problems. Re-recorded audio, real-time voice conversion, very short clips and untested languages are listed as open problems on the research page; the unmeasured ones are worth your time.


Questions from research groups

Can I get score-level output rather than a verdict label?

The right thing to ask for, and not yet defined. Verdict labels are lossy for evaluation because they hide where the threshold sits, and a vendor-chosen threshold is a modelling decision made on your behalf. The terms for score-level access are outstanding rather than refused.

Will results change under my feet during a study?

They can. The engine is retrained and rates re-measured after every retrain, which is why a version string is attached to every report. A study spanning a version boundary without recording which clips were scored under which engine is not reproducible, and version pinning for research use is among the terms still to be defined.

Has anyone independent verified the accuracy claims?

No. No third-party audit, no external replication of the benchmark, no certification of any kind. The claims are self-reported and the method is published so they can be checked. Until somebody outside checks them, read them as a vendor’s own measurement.

I have a clip your engine got wrong. Is that useful?

More useful than almost anything else. Send it with the generation details if synthetic, or the recording chain if genuine, to coverage@aivoicedetctor.com. If you are submitting participant speech collected under ethics approval, read privacy and data retention against your consent terms first.


Start with the parts built to be attacked

The dataset composition and the open problems list are where the weaknesses are written down. Read those before the accuracy page and you will know what the numbers are worth when they arrive.

Reviewed