Truthring
Coverage · licensed voice

Detecting WellSaid Labs

There is a real person behind a WellSaid voice. The company’s stated model is to build its voices with voice actors who are engaged and compensated for it [VERIFY: verify — current terms]. The audio is still generated, and Truthring still identifies it as generated in [VERIFY]% of clean test clips. Holding both of those facts at once is the whole subject of this page.

Detection answers one question and it is not the interesting one. We can tell you a clip was manufactured rather than captured. We cannot tell you whether manufacturing it was allowed. Those are separate enquiries, they are answered by different kinds of evidence, and conflating them is the single most common error we see people make with a report.


Two questions that get confused

Ask what happened to a piece of audio and you are really asking several things at once. It helps to pull them apart, because only one of them is an acoustic question.

The questionWhere the answer livesCan a detector reach it?
Was this audio generated or recorded?In the signalYes — this is the product
Which system generated it?In the signal, weakly and perishablyOften, and the answer decays between retrains
Whose voice is it modelled on?In the vendor’s recordsNo
Did that person agree?In a contractNo
Were they paid, and are they still being paid?In accounts and licence termsNo
Was this particular use inside the licence?In the licenceNo

Four of those six questions are the ones people actually care about, and a detector reaches none of them. That is not a gap we intend to close later. It is a boundary in the physics of the problem.


Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
Exported narration, unmodified[VERIFY]%
Audio extracted from a published video[VERIFY]%
Broadcast-processed and loudness-normalised[VERIFY]%
Named as WellSaid specifically[VERIFY]%
Real recorded voice acting wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips across [VERIFY: n] voices, generated on [VERIFY: date] and held out of training. Composition and method: accuracy and benchmark.


Why this generator is difficult in an unusual way

Most systems on this site are difficult because they are trying to sound like a person. This one is difficult because it was built from a person doing their job well.

Professional voice work has qualities amateur speech does not: consistent distance from the microphone, controlled breath, deliberate pacing, an even noise floor, a studio behind it. Those qualities look, from a signal point of view, faintly artificial even when a human produced them — a booth is an unusually clean acoustic environment, and cleanliness is one of the things our first pass treats as suspicious. Train a system on that material and you inherit its polish.

The practical result is that our two error modes push in opposite directions here. Genuine studio voice acting is our hardest false-positive case, and generated narration modelled on studio voice acting is a harder true-positive case than a hobbyist clone would be. We publish the wrongly-flagged figure above for exactly this reason: on this page it is arguably the more important number.

What each pass contributes here

1

Recording chain

Still useful, but less decisive than elsewhere. The reference material was captured in a treated room, so the absence of room character is less anomalous than it would be on speech that claims to have been recorded in a kitchen.

Weight on this generator: [VERIFY: high / medium / low]

2

Generator signature

Carries the most weight. A catalogue of designed voices rendered through one pipeline produces the kind of repeated regularity a signature is made of, which is also why attribution here is comparatively durable between our retraining runs.

Weight on this generator: [VERIFY: high / medium / low]

3

Prosody under stress

Contributes little. Scripted narration contains no interruptions, no restarts, no laughter, no sentence abandoned halfway. There is nothing at the edges for this pass to look at, because narration has no edges.

Weight on this generator: [VERIFY: high / medium / low]

Where this fails

  • Broadcast processing. Compression, de-essing and loudness normalisation applied in post move both real and generated speech toward the same place. Submit a pre-master if you have one.
  • Genuine studio recordings. Our highest false-positive risk on the whole site sits here. A booth recording of a professional reader is the human speech most likely to be wrongly flagged.
  • Short promotional cuts. Advertising audio is often a few seconds long, which starves the second and third passes.
  • A model newer than our last retrain. The verdict may hold while attribution degrades to unknown generator. Attribution always decays first.

The general list: where Truthring is wrong.

The asymmetry, and one place it cuts against an actor. Likely synthetic is the stronger verdict because positive evidence is required to reach it. Likely human is weaker, because it can also mean the evidence was destroyed in transit. A performer trying to show that a clip attributed to their session was in fact generated has the stronger direction of the test available to them. A performer trying to show the opposite — that a disputed clip really was their own recorded voice — is asking for the weak verdict, and should not rest a claim on it.


Questions

If the actor consented, is it still synthetic?

Yes. Consent changes whether a use was legitimate; it does not change how the audio was made. Both statements can be true of the same clip and usually are.

Can you tell me whether a voice was licensed?

No, and neither can anyone else by listening. Permission lives in paperwork. Any product claiming to hear the difference between authorised and unauthorised synthesis is describing something that does not exist.

I am a voice actor. Is this useful to me?

Up to a point. We can establish that a clip was generated and often name the system. We cannot confirm the voice is modelled on yours — speaker identification is a different task and not one we perform.

Why does attribution matter if the use was fine?

Because it can contradict a story. Audio presented as an original session that analyses as rendered output is a finding regardless of whether synthesis itself was permitted.

Our agency’s voiceover came back synthetic. Should we be worried?

Probably not. A large share of professional narration is now produced this way under licence. Ask what the contract said, not what the waveform says.

Signature last retested [VERIFY: date] against the WellSaid voice set current at [VERIFY: date]. Rates on this page are re-measured monthly and change when the vendor ships.

Reviewed