Truthring
Coverage · expressive speech

Hume, and the end of “it sounded real”

Most people decide whether a voice is genuine by how it makes them feel. Hume is a system built to make speech emotionally expressive — to sound warm, hesitant, amused or distressed on purpose. That is the exact cue human judgement rests on, which is why this page is less about our detection rate and more about a habit you need to give up.

Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
Studio-quality clip[VERIFY]%
Voice note[VERIFY]%
Recorded phone call[VERIFY]%
Highly expressive delivery[VERIFY]%
Named as this system specifically[VERIFY]%

Measured on [VERIFY: n] clips across [VERIFY: n] emotional settings, generated on [VERIFY: date] and held out of training. Dataset composition: research.


What this system is for

Hume builds speech interfaces that respond expressively — voice agents intended to sound as though they are attending to the person rather than reading at them. Almost all of that is ordinary product work: support lines, companions, accessibility, interactive characters. Nothing about the technology is illegitimate.

What matters for detection is the design goal. A system optimised for emotional plausibility is optimised against the listener’s intuition, and the listener’s intuition is the defence most people are actually relying on when they decide whether a call is real.

The habit to give up. “It sounded like her, and she sounded frightened” is not evidence. Distress is generable. So is warmth, hesitation, a catch in the throat, and the particular flatness of someone in shock. If a caller’s emotional state is doing the persuading, that is the part you should trust least, not most.


Why this does not make detection proportionally harder

There is a useful asymmetry here. Emotional realism is a listener-facing property. Our analysis is not a listener.

None of the three passes asks whether a voice sounds sincere. The first asks whether the audio carries the signature of having passed through a real microphone in a real room, or of having been rendered straight to a file. The second looks for the fine regularities a particular synthesis system leaves behind. Only the third touches prosody, and it examines behaviour at conversational edges — interruptions, restarts, abandoned sentences — rather than emotional plausibility.

So a system that gets dramatically better at moving a human listener may move our score very little. That is the honest version of the good news, and it comes with an equally honest caveat: expressive systems do narrow the third pass, and for this generator the weight shifts toward the first two.

1

Recording chain

Unaffected by expressiveness. A generated file is a generated file however feelingly it reads.

Weight on this generator: [VERIFY: high / medium / low]

2

Generator signature

Carries more of the load here than on a plainer system. Also the pass that decays fastest when a vendor ships.

Weight on this generator: [VERIFY: high / medium / low]

3

Prosody under stress

Expected to be under more pressure here than on most generators we cover — a design expectation, not a measurement. Expressive synthesis is trained on precisely the messy, emotionally loaded speech this pass used to find distinctive.

Weight on this generator: [VERIFY: high / medium / low]


Where we fail on this system

  • Short, highly expressive clips. A four-second sob or a single urgent sentence gives the first pass little and the third pass almost nothing.
  • Live conversational deployment. Speech generated turn-by-turn inside a real call, mixed with real network conditions, is harder than a rendered file. Our rate on that is [VERIFY] and it is lower.
  • Audio re-recorded through a speaker. The universal defeat — it hands the file a genuine recording chain and removes our strongest signal.
  • A model newer than our last retrain. The verdict may hold while attribution drops to unknown generator.

The full list: where Truthring is wrong.

What to do instead of listening harder

If a call is emotionally compelling and asking you to act quickly, the emotional content is the least reliable thing in it. Hang up and call the person back on a number you already have. A generated voice, however moving, cannot answer a phone you dialled.

Then, if you still want to know what the clip was, submit it — but understand which question you are asking. Detection tells you afterwards what a recording was. The callback tells you now whether to act.


Questions

Can emotionally expressive AI speech be detected?

Yes. Truthring identifies this system’s output as synthetic in a share of clean test clips that has not been measured yet. Emotional expression is aimed at a listener, and the analysis is not a listener — so a system that gets better at sounding sincere does not get harder for us at the same rate it gets harder for you.

If a voice sounds emotional, does that mean it is real?

No. It was never strong evidence and it is now close to worthless. Distress, warmth, hesitation and urgency are all generable. If a caller’s emotional state is what is persuading you, treat that as the weakest part of the call.

Does expressive synthesis defeat your prosody analysis?

It puts real pressure on it. That pass looks at behaviour at conversational edges rather than at emotional plausibility, but a system trained on expressive speech narrows the gap. For this generator the weight shifts toward the recording-chain and signature passes, and we would rather say so than imply the third pass is untouched.

Is it wrong to build voices like this?

No, and we are not going to pretend otherwise. Expressive speech has obvious value in accessibility and in interfaces people have to use for a long time. The problem is not the technology; it is that a great many defences still rest on an assumption — that a voice which moves you is a voice that belongs to someone — which stopped being true.

Signature last retested [VERIFY: date] against version [VERIFY]. Rates re-measured monthly and change when the vendor ships.

Reviewed