Truthring
Coverage · text to speech

Detecting Amazon Polly

Polly is infrastructure. It reads out platform changes, phone menus, delivery updates, article pages and screen-reader text, and it has been doing so quietly for long enough that most people have heard it today without noticing. Truthring identifies its output as synthetic in [VERIFY]% of clean test clips and names the engine in [VERIFY]% of those. The harder thing this page has to tell you is that on this generator, a synthetic verdict is usually not news.

Start with the base rate. If you are checking a phone menu, an announcement or an article read aloud, the honest prior is that it was synthesised. Confirming it moves you almost nowhere. The question worth asking about a Polly clip is not was this generated but did anyone claim it was not.

Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
Rendered file, unmodified[VERIFY]%
Announcement captured over a public address system[VERIFY]%
Phone menu recorded from a call[VERIFY]%
Named as Amazon Polly specifically[VERIFY]%
Real human speech wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips across [VERIFY: n] stock voices, generated on [VERIFY: date] and held out of training. Composition and method: accuracy and benchmark.


What Amazon Polly is, and where you meet it

Polly is Amazon Web Services’ text-to-speech service. It converts written text into spoken audio in a catalogue of stock voices. It does not exist to reproduce a named individual’s voice, and that single fact shapes everything on this page.

Because it is sold as a component rather than a creative tool, it tends to be embedded rather than performed. It appears inside interactive voice systems, transport and building announcements, accessibility readers, warehouse and dispatch equipment, e-learning modules, and the “listen to this article” button on news sites. In almost none of those settings is anybody pretending a person is speaking. The synthesis is the point, and everybody involved knows it.

That makes Polly the clearest example of a category we think detection products describe badly: speech that is synthetic, ubiquitous and entirely ordinary.

Why a synthetic finding here is weak information

A detector is only as informative as the uncertainty it removes. If you hand us a clip from a category where synthesis is already the norm, we can only tell you what you had good reason to assume. That is not a failure of the analysis. It is a property of the question.

The mistake we want to head off is the one where a synthetic verdict is read as an accusation. Somebody checks a company’s automated callback message, gets likely synthetic, and treats it as evidence of bad faith. It is not. It is evidence that a phone system used a phone system’s voice.

Where a Polly finding does carry weight is where it contradicts a claim. If a file is circulated as a recording of an official reading a statement aloud, and the analysis places it as stock text-to-speech, the two accounts cannot both be true — and that contradiction is a real finding, independent of whether synthesis itself was improper. The same applies when a clip is offered as proof that a particular person said something, and no particular person is present in the audio at all.


How the analysis behaves on stock speech

Three passes run on every submission. The balance between them is different here than on a cloning system, because there is no specific human being imitated and therefore no attempt to reproduce a particular person’s habits.

1

Recording chain

Rendered speech is written straight to a file. There is no room, no microphone, no body producing the sound, and by default no attempt to simulate any of them. On unmodified Polly output this pass is usually decisive on its own.

Weight on this generator: [VERIFY: high / medium / low]

2

Generator signature

Stock catalogues are a closed set. The same handful of voices are produced by the same pipeline millions of times, which makes their regularities comparatively easy to characterise and comparatively stable between our retraining runs. Attribution on this generator ages more slowly than it does on systems that ship frequently.

Weight on this generator: [VERIFY: high / medium / low]

3

Prosody under stress

Contributes least here. Infrastructure speech is short, declarative and never interrupted — a platform number, a menu option, a delivery window. There is no conversational messiness for this pass to examine, so on typical clips it has little to say.

Weight on this generator: [VERIFY: high / medium / low]

Where this fails

  • Announcements captured in the room. A station announcement played through a loudspeaker and recorded on a phone acquires a genuine recording chain — reverberation, crowd noise, an actual microphone. The strongest pass is neutralised by ordinary circumstance rather than by anyone trying.
  • Telephone-band audio. Interactive voice systems are usually heard over a narrowband, heavily compressed channel that discards the fine structure the analysis reads. See what happens to a clip on a phone call.
  • Very short utterances. Two or three words gives the second and third passes almost nothing to work with. Submit the longest continuous stretch you have, not the most quotable fragment.
  • Spliced files. A recording that alternates synthesised prompts with a human caller will return a mixed and lower-confidence result. Cut the segment you actually care about and submit that.

General failure modes that apply regardless of generator: where Truthring is wrong.

The asymmetry, restated for this page. Likely synthetic is the stronger verdict because it requires positive evidence to be present. Likely human is weaker, because it can also mean the evidence was destroyed in transit. On infrastructure audio — loudspeakers, phone lines, re-recordings — evidence is destroyed in transit routinely, so a human-leaning result on this material deserves very little of your confidence.


Questions

Is it worth checking whether a public announcement was synthetic?

Rarely on its own. Recorded announcements, phone menus and reading tools have used synthesised speech for years. A synthetic verdict confirms the expected. It becomes useful only when someone has claimed the recording captured a live human being.

Can you tell Polly apart from other stock engines?

Often. A fixed catalogue rendered by a consistent pipeline leaves a steadier signature than a cloning system does. Attribution still trails detection, and a clip we cannot place reads unknown generator rather than a nearest guess.

My phone menu recording came back inconclusive. Why?

Because the channel removed the evidence before it reached us. Narrowband telephony strips exactly the detail the first two passes depend on. That is a limitation of the audio, not a judgement about it.

Someone used a stock voice to impersonate an official body. Does this help?

As a supporting detail. The deception is in the claimed authority, not the synthesis. What we can supply is the contradiction between “this is a person speaking” and an analysis that places the audio as text-to-speech.

Does a synthetic verdict imply wrongdoing?

No, and on this generator it almost never does. Detection describes how audio was made. It has nothing to say about whether making it that way was appropriate, disclosed or agreed.

Signature last retested [VERIFY: date] against Amazon Polly voice set [VERIFY: which voices]. Rates on this page are re-measured monthly and change when the vendor ships.

Reviewed