The voice you stopped noticing
Google’s cloud text-to-speech, including the WaveNet family of voices, has been in production for years. It reads out platform announcements, navigation instructions, phone menus and accessibility text at a volume that makes it one of the largest single sources of synthetic speech most people hear. Nobody is fooled by it, because nobody was meant to be.
| Audio | Reads as synthetic | Named to this family |
|---|---|---|
| Rendered file, full bandwidth | [VERIFY]% | [VERIFY]% |
| Played over a public address system and recorded | [VERIFY]% | [VERIFY]% |
| Telephone menu, recorded at the handset | [VERIFY]% | [VERIFY]% |
| Embedded in a compressed video | [VERIFY]% | [VERIFY]% |
Measured on [VERIFY: n] clips across [VERIFY: n] voices and [VERIFY: n] delivery paths, generated on [VERIFY: date] and held out of training. Method: accuracy.
Why an older family is an easier target
Detection of a long-deployed synthesis family is a fundamentally different problem from detection of a system that shipped last month, and it is easier in three ways at once.
The first is time. A signature is learned from examples, and a family that has been in production for years has generated an enormous quantity of observable audio in an enormous variety of conditions. Our training set has seen it played through car speakers, station announcements, hold music systems and bad phone lines. Nothing about it surprises us any more.
The second is stability. Attribution decays when a vendor ships a new model — that is the single largest source of drift in this whole product. A family whose behaviour has been broadly consistent for a long stretch gives attribution time to settle, which is why the naming figures in the table above are unusually close to the verdict figures. On a current cloning system the gap between those two columns is typically much wider.
The third is intent. Systems of this generation were built to be clear, consistent and pleasant to listen to. They were not built to convince a listener that a specific human being sat in a specific room in front of a specific microphone. The physical evidence of a recording session is absent because reproducing it was never part of the brief. Current cloning systems, whose entire purpose is to pass as a particular person, work much harder at precisely that.
What changed between then and now
| Cue | Older synthesis families | Current cloning systems |
|---|---|---|
| Room and microphone evidence | Consistently absent; a strong and reliable cue | Sometimes simulated; a weaker cue than it was |
| Prosody at the edges of speech | Even and regular by design | Deliberately varied; trained on expressive material |
| Breath and hesitation | Sparse or formulaic | Reproduced, and often reproduced well |
| Stability of the signature | Slow-moving, so attribution holds | Moves with each release; attribution decays quickly |
| What defeats detection | Channel damage — codecs, public address systems, noise | Channel damage, plus the sophistication of the generator itself |
The common column is the last one. Every generation of synthesis, old or new, is protected by the same thing: a lossy channel between the file and us.
Old audio, checked late
A recurring use of this page is archival. Someone finds a recording from several years ago — a voicemail, a captured announcement, an audio file attached to an old case — and wants to know what produced it. This is one of the few situations where our position genuinely improves with time, because we now hold a mature picture of what the systems of that period sounded like and they have stopped changing.
Two cautions apply. Old files have usually been through more handling than new ones: converted between formats, moved between systems, re-encoded by whatever archived them. Each of those hops removes evidence, and a five-year-old file has had five years of opportunities. And the storage format matters more than its age — a well-kept uncompressed archive from years ago is a far better subject than a recent file that has been through three messaging apps.
The other honest note: an old file that reads synthetic is very often synthetic for a completely mundane reason. Automated announcements, hold messages, reminder calls and accessibility output have been generated for years. The interesting archival result is not this was synthetic but this was synthetic and it was presented as a person speaking, and only the surrounding record establishes the second half.
The asymmetry holds here too. Likely synthetic is the stronger verdict — it is reached only when positive evidence is present, and on a stable older family that evidence is comparatively easy to find. Likely human remains the weaker answer: on a recording that has been through a public address system, a phone line, or a decade of format conversions, it may simply mean the evidence did not survive the journey.
Questions
Are older AI voices easier to detect than new ones?
Usually, and by a clear margin on both the verdict and the attribution. The family is stable, we have observed a great deal of it, and it was never engineered to imitate the physical evidence of a real recording.
Why is your naming rate so much closer to your detection rate here?
Because attribution decays with vendor releases, and this family has moved slowly. On a fast-moving cloning product, attribution can sit far below the verdict figure and drop again the week a new model ships.
My phone company’s menu reads as synthetic. Is that a problem?
No. Phone menus, announcements, navigation prompts and screen readers have been synthetic for years and are meant to be. A verdict here is a description of production, not a finding of anything.
Can you check a recording from several years ago?
Yes, and older systems are among our better subjects. Send the least-handled copy you can find — the original file rather than something re-exported for the purpose — and tell us roughly when it was made, which narrows what we compare it against.
Does the sheer amount of synthetic speech around us make detection pointless?
It changes what detection is for. When most synthetic speech was suspicious, the verdict was nearly the whole answer. Now it is one input, and the judgement lives in the context: who sent this, through what channel, asking for what.
Signature last retested [VERIFY: date] against Google Cloud text-to-speech voice generation [VERIFY: verify]. Rates are re-measured monthly; on this family they move slowly.
Reviewed