Truthring
Coverage · open-weight models

Detecting Tortoise TTS

Tortoise TTS is an open-source speech synthesis project that takes the unusual position of preferring output quality to generation speed — a choice announced in its name. People often read that slowness as a kind of built-in restraint. It is worth being precise about who it restrains, because the answer is: not the person you are worried about.

Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
High-effort render, clean file[VERIFY]%
Fast preset render, clean file[VERIFY]%
Recorded phone call[VERIFY]%
Named as Tortoise specifically[VERIFY]%
Real human speech wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips generated at [VERIFY: which] quality settings and held out of training. Method: how it works.


Who is actually slowed down

Generation time is a cost, and costs deter. The question is which behaviours they deter. A campaign that needs ten thousand distinct calls is genuinely constrained by minutes per clip; the arithmetic stops working and the operator moves to a faster tool. That is a real effect and it is worth something.

But the incidents that do the most damage per occurrence are not volume attacks. They are single clips aimed at one person: a finance officer who receives a voice note from someone who sounds like the chief executive, a parent who takes a call in their child’s voice, a journalist sent a recording of a source saying something they never said. Each of those needs one file. An attacker preparing one file has an evening, a weekend, as long as they like. They will happily wait for a better render, and the wait improves their result.

So the honest summary is that slow generation shifts the mix of abuse rather than reducing it. It filters out the noisy, low-value, high-volume attempts and leaves the patient, targeted ones untouched.

Attack shapeConstrained by generation time?What actually constrains it
Mass robocalling in a cloned voiceYes, substantiallyCost per call, carrier filtering, telephony access
One voice note to one finance approverNoPayment controls and callback procedure
A fabricated recording of a public figureNoProvenance checks before publication
Voice-based account recovery bypassRarelyWhether voice is treated as a factor at all

The trade-off is asymmetric, and it runs against defenders

Someone generating a clip runs the model once. Someone screening clips runs an analysis on everything that arrives — every voice note into a support queue, every recording attached to a claim, every submission to a newsroom tip line. The attacker can spend an hour of compute on one file. The defender has to spend far less than that on each of thousands, or the process does not happen at all.

This is why we publish detection rates by condition rather than a single headline figure. A number produced by an analysis nobody could afford to run at volume would be a marketing claim rather than an operational one. The rates on this page are measured under settings we would actually apply to a queue.

There is a second, more useful asymmetry available to defenders, and it is procedural rather than computational. The attacker must get everything right — the voice, the context, the pretext, the pressure to act now. The defender only has to break one link, and the cheapest link to break is almost never the audio. It is the callback.

Does better-sounding output detect worse?

Not in the way people expect. Perceived quality and detectability are related, but loosely. A render that a listener rates as excellent has been optimised to satisfy a listener — natural rhythm, plausible emphasis, no obvious seams. None of that is the same as reproducing the physical history of a recording, which is what the first pass of our analysis reads.

In practice we see the effect in a specific place: high-effort renders tend to be more internally consistent than fast ones, and consistency of the wrong kind is itself informative. Real speech drifts. A person’s distance from the microphone changes, their voice tires, the room does something unhelpful. Audio that is uniformly excellent from first syllable to last is not what a recording usually looks like.


Where this fails

  • Deliberate degradation. Anyone patient enough to run a slow generator is patient enough to add noise, room reverb and a codec pass afterwards. That is the strongest defeat available and it costs nothing but time, which is the resource this attacker already has.
  • Unfamiliar forks. Open projects accumulate variants. A fork we have not tested can still read as synthetic while attribution falls to unknown generator.
  • Short clips. Under [VERIFY: n] seconds the generator signature carries little weight and the verdict leans on capture history alone.
  • Compressed delivery. A phone network or a messaging app removes much of what the analysis reads — see limitations.

The asymmetry worth understanding. Likely synthetic is the stronger verdict because it requires positive evidence to be present. Likely human is weaker: it can mean the clip is genuine, or that the evidence never survived the journey. Against a patient attacker who has deliberately degraded a file, a likely human result should be read as an absence of findings, not as a clearance.


If you are assessing one high-stakes clip

The advice changes when the volume is one. Submit the longest continuous stretch of speech you hold rather than the most incriminating sentence — length is the input that most reliably raises confidence. Submit the original file, not a copy that has been forwarded through an app. Note how the clip reached you, because the delivery path usually tells you more than the audio does. And treat the confidence figure as the result: [VERIFY]% and 51% both read likely synthetic and mean entirely different things.


Questions

Is a slow generator less dangerous?

Less useful for volume, no less useful for a single targeted clip. Since targeted clips cause most of the per-incident harm, treating slowness as a safety property is a mistake.

Does Truthring take longer on high-quality renders?

No. Our analysis cost depends on clip length and format, not on how much effort went into producing the audio. The rates published here are measured under conditions we can sustain at volume.

Will the verdict name Tortoise?

When the clip matches the signature we hold, yes. Forks and modified checkpoints frequently detect as synthetic while attributing to unknown generator, and we prefer that blank to a guess.

Does the quality setting used affect the result?

Somewhat, which is why the two render conditions are listed separately above rather than averaged into one figure.

Can I use a result to accuse someone?

No. A verdict is a probability with a published method and a published error rate. It supports a decision to verify further; it does not settle a question about a person on its own.

Signature last retested [VERIFY: date] against build [VERIFY: which] at [VERIFY: which] quality presets. Rates are re-measured monthly.

Reviewed