Truthring
Benchmark · 2026 · engine v2.4

Truthring benchmark 2026

An open comparison of AI voice detectors on telephone-quality audio. The dataset is described, the metrics are standard, and results are published whether or not they favour us.

Read this first. We built this benchmark and we are in it. That is a conflict of interest, and the only honest response is to publish the method fully enough that someone else can run it and check. The dataset description and scoring code are linked below.


Results — telephone-quality audio

DetectorPrecisionRecallF1False positiveAUROC
Truthring v2.4[VERIFY][VERIFY][VERIFY][VERIFY][VERIFY]
aivoicedetector.com[VERIFY][VERIFY][VERIFY][VERIFY][VERIFY]
ElevenLabs Classifier[VERIFY][VERIFY][VERIFY][VERIFY][VERIFY]
Resemble Detect[VERIFY][VERIFY][VERIFY][VERIFY][VERIFY]
voiceaichecker.com[VERIFY][VERIFY][VERIFY][VERIFY][VERIFY]

Tested [VERIFY: date] against each product's publicly available free or trial tier. Vendors were not contacted in advance. Where a detector returns a probability rather than a verdict, the threshold used is stated in the method below.


Results — clean studio audio

DetectorPrecisionRecallF1False positiveAUROC
Truthring v2.4[VERIFY][VERIFY][VERIFY][VERIFY][VERIFY]
aivoicedetector.com[VERIFY][VERIFY][VERIFY][VERIFY][VERIFY]
ElevenLabs Classifier[VERIFY][VERIFY][VERIFY][VERIFY][VERIFY]
Resemble Detect[VERIFY][VERIFY][VERIFY][VERIFY][VERIFY]

[VERIFY: If a competitor beats you here, say so in one plain sentence directly beneath this table.] A benchmark that only ever favours its author is not a benchmark.


The dataset

ClassSamplesSource
Synthetic — commercial TTS[VERIFY][VERIFY: which systems]
Synthetic — voice cloning[VERIFY][VERIFY]
Synthetic — open source[VERIFY][VERIFY]
Genuine human — studio[VERIFY][VERIFY: source and consent basis]
Genuine human — telephone[VERIFY][VERIFY]
Total[VERIFY]
  • Degradation applied: [VERIFY: codecs, bitrates, sample rates, noise profiles]
  • Held out from training: [VERIFY: confirm explicitly]
  • Availability: [VERIFY: whether the set is shared, with whom, under what license]

Full description at research/dataset.


Method

  • Every detector received the identical file for each sample — no re-encoding between tools.
  • Where a detector returns a probability, the decision threshold used was [VERIFY], chosen [VERIFY: how].
  • Where a detector returns a third "uncertain" state, those results are reported separately rather than forced into a binary. [VERIFY: state how you handled this — it materially affects the numbers.]
  • Each sample was submitted once. No retries, no cherry-picking.
  • Scoring code: [VERIFY: repository link, or say it is available on request]

What this benchmark does not show

It measures accuracy on one dataset at one point in time. It does not measure speed, price, API reliability, support, or how each product behaves on audio unlike ours. Detectors are retrained frequently — a result from August 2026 may not hold in December.

Where a competitor's product is aimed at a different use case, comparing them on ours is not a fair test of what they built. We have noted those cases in the tables above.


Reproduce it, or correct it

If you are one of the vendors listed and believe we have tested your product incorrectly — wrong tier, wrong threshold, wrong handling of your uncertain state — write to us and we will re-run it and publish the correction alongside the original.

We do not quietly edit results. benchmark@aivoicedetctor.com


Authorship

[VERIFY: Written by — real name, real role.] Benchmark run [VERIFY: date]. Next scheduled [VERIFY].

Reviewed