Detecting WellSaid Labs
There is a real person behind a WellSaid voice. The company’s stated model is to build its voices with voice actors who are engaged and compensated for it [VERIFY: verify — current terms]. The audio is still generated, and Truthring still identifies it as generated in [VERIFY]% of clean test clips. Holding both of those facts at once is the whole subject of this page.
Detection answers one question and it is not the interesting one. We can tell you a clip was manufactured rather than captured. We cannot tell you whether manufacturing it was allowed. Those are separate enquiries, they are answered by different kinds of evidence, and conflating them is the single most common error we see people make with a report.
Two questions that get confused
Ask what happened to a piece of audio and you are really asking several things at once. It helps to pull them apart, because only one of them is an acoustic question.
| The question | Where the answer lives | Can a detector reach it? |
|---|---|---|
| Was this audio generated or recorded? | In the signal | Yes — this is the product |
| Which system generated it? | In the signal, weakly and perishably | Often, and the answer decays between retrains |
| Whose voice is it modelled on? | In the vendor’s records | No |
| Did that person agree? | In a contract | No |
| Were they paid, and are they still being paid? | In accounts and licence terms | No |
| Was this particular use inside the licence? | In the licence | No |
Four of those six questions are the ones people actually care about, and a detector reaches none of them. That is not a gap we intend to close later. It is a boundary in the physics of the problem.
Measured on [VERIFY: n] clips across [VERIFY: n] voices, generated on [VERIFY: date] and held out of training. Composition and method: accuracy and benchmark.
Why this generator is difficult in an unusual way
Most systems on this site are difficult because they are trying to sound like a person. This one is difficult because it was built from a person doing their job well.
Professional voice work has qualities amateur speech does not: consistent distance from the microphone, controlled breath, deliberate pacing, an even noise floor, a studio behind it. Those qualities look, from a signal point of view, faintly artificial even when a human produced them — a booth is an unusually clean acoustic environment, and cleanliness is one of the things our first pass treats as suspicious. Train a system on that material and you inherit its polish.
The practical result is that our two error modes push in opposite directions here. Genuine studio voice acting is our hardest false-positive case, and generated narration modelled on studio voice acting is a harder true-positive case than a hobbyist clone would be. We publish the wrongly-flagged figure above for exactly this reason: on this page it is arguably the more important number.
What each pass contributes here
Recording chain
Still useful, but less decisive than elsewhere. The reference material was captured in a treated room, so the absence of room character is less anomalous than it would be on speech that claims to have been recorded in a kitchen.
Generator signature
Carries the most weight. A catalogue of designed voices rendered through one pipeline produces the kind of repeated regularity a signature is made of, which is also why attribution here is comparatively durable between our retraining runs.
Prosody under stress
Contributes little. Scripted narration contains no interruptions, no restarts, no laughter, no sentence abandoned halfway. There is nothing at the edges for this pass to look at, because narration has no edges.
Where this fails
- Broadcast processing. Compression, de-essing and loudness normalisation applied in post move both real and generated speech toward the same place. Submit a pre-master if you have one.
- Genuine studio recordings. Our highest false-positive risk on the whole site sits here. A booth recording of a professional reader is the human speech most likely to be wrongly flagged.
- Short promotional cuts. Advertising audio is often a few seconds long, which starves the second and third passes.
- A model newer than our last retrain. The verdict may hold while attribution degrades to unknown generator. Attribution always decays first.
The general list: where Truthring is wrong.
The asymmetry, and one place it cuts against an actor. Likely synthetic is the stronger verdict because positive evidence is required to reach it. Likely human is weaker, because it can also mean the evidence was destroyed in transit. A performer trying to show that a clip attributed to their session was in fact generated has the stronger direction of the test available to them. A performer trying to show the opposite — that a disputed clip really was their own recorded voice — is asking for the weak verdict, and should not rest a claim on it.
Questions
If the actor consented, is it still synthetic?
Yes. Consent changes whether a use was legitimate; it does not change how the audio was made. Both statements can be true of the same clip and usually are.
Can you tell me whether a voice was licensed?
No, and neither can anyone else by listening. Permission lives in paperwork. Any product claiming to hear the difference between authorised and unauthorised synthesis is describing something that does not exist.
I am a voice actor. Is this useful to me?
Up to a point. We can establish that a clip was generated and often name the system. We cannot confirm the voice is modelled on yours — speaker identification is a different task and not one we perform.
Why does attribution matter if the use was fine?
Because it can contradict a story. Audio presented as an original session that analyses as rendered output is a finding regardless of whether synthesis itself was permitted.
Our agency’s voiceover came back synthetic. Should we be worried?
Probably not. A large share of professional narration is now produced this way under licence. Ask what the contract said, not what the waveform says.
Signature last retested [VERIFY: date] against the WellSaid voice set current at [VERIFY: date]. Rates on this page are re-measured monthly and change when the vendor ships.
Reviewed