Truthring
Coverage · assistive speech

Detecting Speechify

Speechify is a commercial text-to-speech product used mainly as a reading aid — documents, articles, textbooks and email turned into audio so they can be listened to rather than read. Truthring identifies its output as synthetic in [VERIFY]% of clean test clips. Before the numbers, though, this page has to say something about what that finding is for, because on this generator the risk is not that we miss a clip. It is that somebody misuses a hit.

This is synthetic speech used for a person, not against one. Almost every Speechify clip that exists was produced so that a human being could get at written material more easily. A detector pointed at that population is not finding fraud. It is finding people who read differently, and a “synthetic” label attached to them is a harm we would rather not have built.


Who is actually on the other end of these clips

The people we hear from about reading tools are dyslexic students working through a set text, commuters getting through a report, people with visual impairments who have used screen readers their whole lives, and people recovering from injury or living with a condition that makes sustained reading exhausting. Some use synthesised speech to take material in. A smaller number use it to put material out — recording a spoken assignment, a voice message or a presentation using generated audio rather than their own voice, because their own voice is unreliable, effortful or hard for others to understand.

That second group is the one this page is written for, because they are the group a detector can hurt. When they submit spoken work, an automated check will find exactly what it is designed to find. The audio was generated. That is true, and on its own it is close to meaningless.

What the verdict is and is not evidence of

Detection answers one question: was this audio manufactured or captured? It cannot answer any of the questions that would actually matter in a dispute. It does not know who wrote the words. It does not know whether an accommodation was agreed. It does not know whether the speaker could have spoken the sentence themselves and chose not to, or could not and had no choice. It does not know whether anyone was told.

Written down like that the limitation is obvious. In practice it gets lost, because a confident-looking percentage next to the word synthetic reads like an accusation even when the report says nothing of the kind. If you are the person receiving such a report about somebody else, the useful next step is almost never another test. It is a conversation.

Reasonable use of a result

  • Checking a single clip you have a specific reason to doubt
  • Establishing whether a voice message attributed to a named person was generated
  • Resolving a contradiction between what a file claims and what it contains
  • Adding one data point to an inquiry that has other evidence in it

Use we ask you not to make

  • Screening a whole class, cohort or workforce to see who comes back synthetic
  • Treating a synthetic finding as proof of dishonesty without asking the person
  • Requiring recorded speech in a process where a synthetic result triggers a penalty
  • Recording the outcome anywhere it would function as an undisclosed disability flag

Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
Exported audio, unmodified[VERIFY]%
Voice note sent through a messaging app[VERIFY]%
Played aloud and captured by a device microphone[VERIFY]%
Named as Speechify specifically[VERIFY]%
Real human speech wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips across [VERIFY: n] voices, generated on [VERIFY: date] and held out of training. Composition and method: accuracy and benchmark.

What the analysis has to work with

Reading tools sit awkwardly in our coverage for a technical reason as well as an ethical one. A product of this kind is a front end: it takes text and returns audio, and the voice doing the speaking may come from more than one underlying engine [VERIFY: verify]. Attribution therefore behaves less predictably here than on a system that renders everything through a single pipeline. A clip can be firmly synthetic and still come back as unknown generator, and that is the correct result rather than a shortfall.

The first pass — whether the audio carries the signature of a real microphone in a real room — does most of the work, as it does with any rendered speech. The third pass, which examines how a voice behaves under conversational stress, has little to examine: material read from a page contains no interruptions, no restarts, no abandoned sentences. Long-form reading does give us one advantage the short clips elsewhere on this site do not have, which is duration. A chapter is easier than a sentence.

Where this fails

  • Playback capture. Audio played through a laptop speaker and recorded on a phone gains a genuine recording chain. Common in classroom settings and enough to defeat the strongest pass.
  • Messaging compression. A voice note forwarded through a chat app is re-encoded at every hop, and each hop removes evidence in both directions.
  • Mixed audio. A recording that alternates generated passages with the speaker’s own voice returns a muddled whole-file result. Submit the segment in question, not the file.
  • Atypical human speech. This is the one that matters most here. Speech affected by a neurological condition, a speech difference, a tracheostomy or profound hearing loss can carry prosody our third pass was not built around. We report a false-positive rate of [VERIFY]% overall and we are explicit that it is not evenly distributed across speakers [VERIFY: verify — per-group figures].

The general list: where Truthring is wrong.

The asymmetry. Likely synthetic is the stronger verdict, because reaching it requires positive evidence. Likely human is weaker, because it can also mean the evidence was destroyed in transit. Neither verdict is a statement about a person’s honesty, and on this generator neither should ever be recorded as one.


Questions

Is using a reading tool cheating?

No. It changes the channel through which someone receives text. It does not change who wrote it, who understood it or who is answerable for it.

A student’s spoken submission reads as synthetic. What now?

Ask them. There are more ordinary explanations than dishonest ones, several of which the student may have had no obligation to disclose to you specifically.

Can you name Speechify in the verdict?

Sometimes. Where the underlying voice does not match a signature we hold, the report reads unknown generator. We do not guess at the nearest match.

We want to screen all submitted audio. Can we?

Technically yes, and we would rather you did not. A blanket screen of a population produces a list of the people in it who use assistive technology. That is not a security finding, and building a process around it creates a liability rather than removing one.

Does a human verdict prove someone spoke for themselves?

No. It is the weaker result and it is frequently what a lossy channel looks like. Do not issue it as a clearance.

Signature last retested [VERIFY: date] against the voice set current at [VERIFY: date]. Rates on this page are re-measured monthly and change when the vendor ships.

Reviewed