You have heard it without being told
OpenAI provides speech synthesis that other developers build into their own products. That single fact makes this page different from the rest of the coverage section. Most people who encounter this output do so inside somebody else’s app — a reading assistant, a language tutor, a support widget, a feature inside a piece of software that mentions no speech vendor anywhere in its interface.
Measured on [VERIFY: n] clips across [VERIFY: n] stock voices, generated on [VERIFY: date] and held out of training. Method: accuracy.
Two different claims, sold in one report
Every result we return contains two statements, and they are not equally strong. The first is a verdict: this audio was manufactured rather than captured. The second is an attribution: it was manufactured by a system whose fingerprint we recognise.
The verdict is the durable part. It rests on broad properties — the absence of a physical recording chain, the way prosody behaves at the edges of an utterance — that hold across generators and survive a fair amount of handling. It is the claim we are willing to defend, and it is the one almost everybody actually needs.
Attribution is narrower and more fragile. It rests on small regularities characteristic of one synthesis family, and those regularities are the first thing lost to compression and the first thing invalidated when a vendor ships a new model. Attribution is the part of the report that decays fastest between our retraining runs, which is why this page carries a retest date and why the two figures above differ.
So when a report reads likely synthetic, unknown generator, nothing has gone wrong. That is an honest description of a common situation: enough evidence to say the audio was made, not enough to say by what. We do not fill the gap with a nearest match. A wrong name attached to a real company is a harm we cannot take back, and the person relying on it is worse off than if we had said nothing.
Where the vendor goes missing
| How you met the audio | What you can usually learn |
|---|---|
| An app that reads articles aloud | That it is synthetic. The speech vendor is often disclosed only in the app’s documentation, if at all. |
| An assistant feature inside other software | That it is synthetic. Vendors are commonly swapped between releases without any user-visible change. |
| A support agent that answered a chat by voice | That it is synthetic, plus whatever the operator discloses. The operator, not the speech vendor, holds the obligation. |
| A clip forwarded with no context | The verdict, and attribution only if the signature is current and the audio is clean. |
A useful habit: separate who generated this from who published this. The second is answerable and usually matters more. The first is often unknowable from the audio and rarely changes what you should do.
Stock voices, and the thing they are not
A stock synthetic voice is not a clone of anybody. It does not claim to be a particular person and nobody is impersonated by it. That places most of this output in a different category from the cloning systems elsewhere in this section: the deception risk is lower, and the disclosure question is the interesting one.
It also creates a specific and underappreciated confusion. People sometimes recognise a stock voice as sounding like someone they know, or like a voice they have heard in another product, and conclude that a clone was made. Widely deployed stock voices turn up in thousands of unrelated places, which is exactly why they sound familiar. Familiarity is not evidence of cloning.
The genuine risk in a widely embedded engine is different and more mundane. Synthetic speech that people encounter constantly stops registering as synthetic. Someone accustomed to hearing generated voices from legitimate apps all day is measurably less alert to a generated voice arriving with a request for money. Ubiquity does the attacker’s work of normalisation for them.
Which verdict carries weight. Likely synthetic is the stronger result: it requires positive evidence in the file. Likely human is weaker, because it can also mean the evidence was destroyed in transit — and audio from an app has usually been encoded at least once before you ever hear it. Read the confidence figure alongside the label; two results with the same label and very different confidence are very different findings.
Questions
Can this output be detected?
On clean audio, a share of held-out clips that has not been measured yet read as synthetic. Attribution to this specific vendor sits at a rate not yet measured and falls faster than the verdict once audio has been compressed.
My report said synthetic but would not name a generator. Is that a failure?
No, it is the expected outcome for output from a system newer than our last training run, or for audio that has lost detail in transit. The verdict is the product; attribution is the bonus.
How do I find out which speech engine an app uses?
Its documentation, its subprocessor list, or its privacy notice, in roughly that order. This is a question about a company’s supply chain rather than about a waveform, and the audio is a poor way to answer it.
Does a stock voice sounding like my friend mean someone cloned them?
Almost certainly not. Stock voices are deployed in enormous numbers of products and coincidental resemblance is common. A clone question needs more than a resemblance to get started.
If everything is synthetic now, is detection still worth anything?
Yes, but the value shifts. As synthetic speech becomes ordinary, the interesting output is not the label on its own but the label plus context: an unexpected voice, an unexpected channel, an unexpected request. Detection is one input into that judgement, not a substitute for it.
Signature last retested [VERIFY: date] against OpenAI speech model version [VERIFY: verify]. Attribution rates are re-measured monthly and drop whenever a vendor ships before we retrain.
Reviewed