Truthring
Coverage · voice cloning

Detecting ElevenLabs

ElevenLabs is the system behind most cloned voices people actually encounter — the fake CEO on a call, the parent’s voice in a ransom hoax, the narrator of a video that never had a narrator. Truthring identifies its output as synthetic in [VERIFY]% of clean test clips and names it as the generator in [VERIFY]% of those. On audio that has been through a phone network the rate is [VERIFY]%.

Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
Studio-quality clip[VERIFY]%
Voice note (WhatsApp, Telegram)[VERIFY]%
Recorded phone call[VERIFY]%
Named as ElevenLabs specifically[VERIFY]%
Real human speech wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips from [VERIFY: n] distinct cloned voices, generated on [VERIFY: date] and held out of training. Method and dataset composition: accuracy and benchmark.


What ElevenLabs is

ElevenLabs is a commercial speech company whose product does two things that matter here. It reads text aloud in a stock voice, which is unremarkable, and it reproduces a specific person’s voice from a sample of their speech, which is not.

The second capability is why the name appears in fraud reports rather than in product reviews. A clone can be built from material that is already public: a conference talk, a podcast appearance, a video posted to a family account, a voicemail left on a colleague’s phone. The account holder confirms they hold the rights to the voice by ticking a box. Nothing verifies the claim at the moment of upload.

The practical consequence is that the barrier to impersonating someone by voice is no longer skill or equipment. It is a minute of their recorded speech and a subscription.

Why these clips are difficult

Detectors built for the previous generation of speech synthesis looked for the things that made it sound wrong: flat pitch, mechanical timing, phonemes stitched at audible seams. Those cues have largely gone. Current cloning systems reproduce breath, hesitation, the drift of pitch across a sentence, and the particular way a person clips the end of a word.

What has not gone is the fact that the audio was manufactured rather than captured. Real speech is recorded through a physical chain — a room, a body, a microphone — and each stage leaves traces that the generator has no reason to reproduce faithfully, because reproducing them was never the goal. The goal was to sound right to a listener.

That gap between sounds right and was recorded is what Truthring measures. It is also why compression hurts us: a phone codec discards exactly the fine structure the analysis depends on.

What Truthring measures on an ElevenLabs clip

Three passes run on every submission. For this generator, the second carries the most weight.

1

Recording chain

Whether the clip carries the signature of having passed through a real microphone and a real room, or of having been rendered directly to a file. Room reflection, handling noise and the noise floor of an actual sensor are hard to fake convincingly and are usually absent by default.

Weight on this generator: [VERIFY: high / medium / low]

2

Generator signature

Fine-grained regularities left by a specific synthesis system — the part of the analysis that lets a verdict name ElevenLabs rather than say only that something is synthetic. Signatures shift when the vendor ships a new model, which is why this page carries a retest date and why we retrain monthly.

Weight on this generator: [VERIFY: high / medium / low]

3

Prosody under stress

How the voice behaves at the edges: an interruption, a laugh, a word restarted mid-syllable, a sentence that trails off. Cloned voices are trained on clean read speech and are least convincing where real conversation is messiest.

Weight on this generator: [VERIFY: high / medium / low]


Where this fails

Four conditions degrade the result on ElevenLabs audio specifically. We would rather you know them before you rely on a verdict.

  • Clips under [VERIFY: n] seconds. There is not enough material for the second and third passes to contribute. Short clips lean almost entirely on the recording-chain analysis and the confidence figure reflects that.
  • Audio re-recorded through a speaker. Playing a generated clip aloud and capturing it on a phone gives the file a genuine recording chain. This is the single most effective way to defeat the first pass, and it is not difficult.
  • A model newer than our last retrain. The verdict may still be correct and read synthetic, but attribution drops to unknown generator. Attribution is the part that decays fastest.
  • Heavy background noise. Traffic, a crowd, music under speech. Noise masks the fine structure in both directions — it can hide a generated clip and it can make a real one look manufactured.

The full list, including the failure modes that apply to every generator rather than this one: where Truthring is wrong.

The asymmetry worth understanding. A likely synthetic verdict on this generator is the stronger of the two results — it is reached when positive evidence is present. Likely human is weaker, because it can also mean the evidence was destroyed in transit. A quiet clean clip that reads human is meaningful. A compressed phone clip that reads human is close to no information at all.


If you have received a clip you think is a clone

Do the cheap thing first, before you upload anything. Hang up and call the person back on the number you already have stored. A voice clone cannot answer a phone you dialled. Almost every voice fraud that succeeds does so because the target stayed inside the channel the attacker chose.

Then, if you still need to know what the clip was:

  • Keep the original file. Do not forward it through a messaging app before checking it — each hop re-compresses the audio and removes evidence.
  • Submit the longest continuous stretch of speech you have rather than the most incriminating sentence.
  • Read the confidence figure, not just the verdict. [VERIFY]% and 51% are both likely synthetic and they mean very different things.
  • If money, a job, a legal claim or someone’s safety depends on the answer, treat the result as one input and confirm through a channel you control.

Questions

Can ElevenLabs voices be detected?

Yes. On clean recordings Truthring identifies ElevenLabs output as synthetic in a share of test clips that has not been measured yet and names the generator in a rate not yet measured of those. Detection falls to a rate not yet measured once the clip has been through a phone call or a messaging app.

Does someone need my consent to clone my voice?

Legally, in most places, yes. Technically, no. ElevenLabs asks an account holder to confirm they hold rights to the voice they upload; that confirmation is a checkbox rather than a verification, so a clone can be built from a podcast episode or a voicemail without the speaker ever knowing.

How much of my voice does someone need?

Roughly a minute of clear speech is enough for a likeness that survives a phone call. More material makes a better clone, but the threshold for a usable one is low enough that anyone who has spoken publicly should assume it is reachable.

Will the verdict tell me it was ElevenLabs specifically?

When the clip matches a signature we hold, yes, and the report names it. When it does not, the verdict reads unknown generator. We do not guess at the nearest match — a confident wrong attribution is more damaging than an honest blank.

Is a Truthring result usable as evidence?

It is a probability with a stated method and a stated error rate, not proof. That makes it usable as one exhibit among others and unusable on its own. Every result carries a reference code so the analysis can be reproduced and challenged.

Signature last retested [VERIFY: date] against ElevenLabs model version [VERIFY]. Detection rates on this page are re-measured monthly and change when the vendor ships.

Reviewed