Truthring
Coverage · legacy synthesis

Detecting iSpeech

iSpeech is a long-established commercial speech company, and its output belongs to a generation of synthesis that predates the current one. That makes it the easiest material on this site to call correctly — [VERIFY]% on clean clips — and it makes this page an argument rather than a scoreboard. Easy to detect is not the same as unlikely to be encountered, and coverage that only tracks whatever shipped this quarter is coverage with a hole in it.

[VERIFY]% Detected as synthetic on clean audio [VERIFY: n] held-out clips
[VERIFY]% Named as this engine rather than generic synthetic Attribution, same set
[VERIFY]% Still detected after telephone-band compression [VERIFY: codec and bitrate]

What the previous generation left behind

The systems that dominated speech synthesis before the current wave were not attempting to be mistaken for a recording. They were attempting to be intelligible, cheap to run and quick to respond, on hardware that was not powerful. Naturalness was a nice-to-have that lost every argument against latency.

Audio built under those constraints carries the constraints in it. Timing tends to be too even, because it was scheduled rather than felt. Pitch moves in ways that resolve too tidily. Transitions between sounds can be abrupt in a manner no vocal tract manages, since a physical tract cannot teleport between shapes. The noise floor is often perfectly still, or perfectly absent, in a way no microphone in any room achieves.

Every one of those is a positive finding rather than an absence, and positive findings are what a detector is good at. We expect confidence on this material to be both higher and steadier than on a current cloning system — less spread between the best and worst clips in a batch. That is a prediction from what the signals look like, not a measurement, and the first benchmark will confirm it or correct it. [VERIFY: spread not measured]

Easy is not rare

There is a temptation, when a system is superseded, to treat it as finished. Deployed software does not work that way. Speech gets embedded into something once and then keeps running as long as the something does, which in the case of phone systems, building announcements, kiosks, dispatch equipment, alarm panels and long-lived mobile applications can be a very long time indeed. Nobody schedules a replacement for a voice that is still saying the right words.

There is also the archive. A recording made years ago becomes relevant when a dispute arises, not when it was made, and disputes have a habit of arriving late. A clip in a legal matter, an insurance claim, a records request or a journalistic investigation may well be older than the current generation of synthesis entirely. A detector that has quietly stopped covering old engines will produce a confident-sounding likely human on exactly that material, which is the worst possible failure: not an error you can see, but silence where a finding should have been.

Attribution has an odd advantage here too. Signatures on current systems perish, because vendors ship and the pipeline changes underneath us. A signature for an engine that has stopped changing does not perish. It is one of the few places on this site where the retest date on a page is a formality rather than a warning.


Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
Rendered file, unmodified[VERIFY]%
Recorded from a phone system[VERIFY]%
Archived audio of unknown provenance[VERIFY]%
Named as iSpeech specifically[VERIFY]%
Real human speech wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips, generated on [VERIFY: date] and held out of training. Where the engine version is not recoverable from the audio, the clip is grouped as [VERIFY: verify — grouping rule]. Composition and method: accuracy and benchmark.

What the three passes do with old audio

The balance is unusual. Our first pass, which asks whether audio carries the trace of a microphone and a room, finds the answer quickly and unambiguously — nothing was recorded and nothing pretends it was. The second pass, generator signature, is unusually reliable for the reason given above. The third, which examines how a voice holds together under conversational stress, is close to irrelevant: this generation of synthesis does not attempt conversational speech at all, so there is no stress behaviour to compare against.

In other words, the material is easy for the two passes that measure production and useless for the pass that measures performance. That is worth knowing if you are comparing figures between this page and, say, a current cloning system, where the weighting is nearly reversed.

Where this fails

  • Narrowband telephony. Most surviving deployments of older synthesis live inside phone systems, which is also the channel that removes the most evidence. See what a phone call does to a clip.
  • Generational analogue loss. Archived material copied between formats accumulates noise and wow that can imitate the irregularity of real speech.
  • Deliberate roughening. Adding hiss, hum or a room impulse to old synthesis is easy and is the obvious way to attack the first pass on this material.
  • Provenance you cannot establish. We can say a clip was generated. We cannot say when, and nothing in the audio dates it.

The general list: where Truthring is wrong.

The asymmetry. Likely synthetic is the stronger verdict, since positive evidence is needed to reach it — and on this generation that evidence is abundant. Likely human stays weak, because it can also mean the evidence was destroyed in transit, and archived or telephone-carried audio is exactly the material most likely to have lost it. High headline accuracy on this page does not make the human-leaning result any more trustworthy than it is elsewhere.


Questions

Is older synthesis easier to detect?

Yes, and more consistently so. It was never built to pass as a recording, so it leaves artefacts a detector reads directly rather than inferring.

If a person can hear it, why run a check?

Because “it sounded odd to me” is not repeatable and does not survive a challenge, and because a clip that has been through a phone line may not sound odd to anyone at all.

Does an old-sounding voice mean an old recording?

No. Older engines remain available and are sometimes chosen on purpose, including by people who want audio to sound like an automated system rather than a person. Nothing in a waveform carries a date.

Why keep a signature for something superseded?

Because the material has not gone anywhere, and because the signature costs almost nothing to maintain. Coverage that tracks only what is new fails precisely on archives, which is where disputes tend to live.

Can you tell which version produced a clip?

Not reliably, and we do not claim to. The report names the engine family where the evidence supports it and stops there [VERIFY: verify — what the report actually prints].

Signature last retested [VERIFY: date] against iSpeech output rendered on [VERIFY: date]. Rates on this page move less than elsewhere on the site, because the engine behind them is no longer changing.

Reviewed