Truthring
Reference · glossary

Glossary

Thirteen terms that people get wrong in ways that change what they believe a result means. Each one has a page because the distinction it carries is load-bearing; the shorter terms are defined here, in place.

In short

The vocabulary of synthetic speech splits into four questions: what a system does to a voice, what a detector measures, what an attack is called, and what a result actually supports. Most costly mistakes in this field come from answering one of those questions with a tool built for another.

How to read this page. The groups below are ordered by what they are about rather than alphabetically, because the confusions worth fixing are between neighbours. If you are here for one word, the fastest route is usually to read it alongside the term it is most often mistaken for — each page names that neighbour explicitly and links to it.


What the systems do

Four different things can be done to or with a voice, and reporting collapses them into “AI voice” almost every time. The difference decides what an attacker is capable of: a generated clip is fixed once made and cannot answer an unexpected question, while a converted voice is driven by a person who can. It also decides whose consent was involved, which is usually the part that matters legally.

Defined here rather than given a page of their own:

  • Synthetic media. The parent category, taking in generated images, video, music and speech at once. A serviceable label for platform rules and a poor unit of technical work: the evidence that betrays a fabricated face has no counterpart in a voice note, and neither does the expertise.
  • Speech-to-speech. A common synonym for voice conversion. Named for its input rather than its output: speech goes in, speech comes out, and text is never involved.
  • Zero-shot cloning. Building a usable clone from a short reference sample without training a dedicated model for that speaker. It is the reason a public video is now enough material to work from.
  • Vocoder. The component that reconstructs an actual waveform from a compressed intermediate representation of sound. Because reconstruction has to invent detail, it is one of the places generation traces are found.
  • Voiceprint, or speaker embedding. The numerical representation of a person’s voice that a verification system stores at enrolment and compares against later. In most regimes it is biometric data about an identified individual.
  • Speaker identification. Searching a voice against many enrolled speakers to find the best match, rather than checking it against one claimed identity. A harder problem than verification and a much noisier one.

What detection measures

A detector does not listen. It computes properties of a signal and compares them against what generated and captured audio tend to look like, and every one of those properties can be destroyed in transit. Understanding which measurement a claim rests on tells you when to distrust it — particularly on telephone audio, where the most useful evidence is discarded by the connection itself.

Defined here rather than given a page of their own:

  • Spectrogram. A picture of a recording with time along one axis and frequency along the other, brightness showing energy. It reveals a file’s handling clearly and its origin barely at all.
  • Formant. A resonance of the vocal tract that gives a vowel its identity. Formants move continuously in real speech, because a tongue has mass and cannot jump between positions.
  • Codec. The scheme used to compress audio for storage or transmission. Every pass through one discards detail permanently, which is why a forwarded voice note carries less evidence than the original recording did.
  • Artefact. A trace left in a signal by the process that produced or handled it, rather than by the sound itself. Detection looks for artefacts of generation and has to distinguish them from artefacts of compression.
  • Noise floor. The quiet background level a real microphone and a real room always produce. Its continuity across a recording is a forensic signal; digitally flat silence is not something a microphone makes.

What the attacks are called

Attack names describe either a channel or a technique, and mixing the two is expensive. A budget written against the technique buys instruments that examine files once the money has gone. A budget written against the channel changes who may authorise an instruction and on what evidence, which is the only one of the two that also closes the version of the attack in which no model was used at all.

Defined here rather than given a page of their own:

  • Smishing. The same social engineering delivered by text message. Named the same way vishing is, by the channel rather than by the tooling.
  • Pretext. The story that makes a request seem reasonable: the compromised account, the supplier changing bank details, the relative in trouble. Pretexts long predate synthesis and have barely changed.
  • Replay attack. Playing a recording of a genuine speaker into a system that authenticates by voice. The oldest spoof against voice biometrics, and it requires no model at all.
  • Presentation attack detection, or anti-spoofing. The component that decides whether a sample offered to a verification system is a live speaker rather than a recording or a synthesis. It sits alongside the verifier and is measured separately — if a vendor quotes only matching accuracy, spoof resistance has probably not been measured.
  • Cheapfake. Misleading media made with ordinary editing rather than generation: a genuine clip trimmed, slowed or re-captioned. It carries no generation signature, so a synthesis detector will pass it.

What a result means

This is the group where careless language costs the most, because these words are what someone quotes when a decision is being made about a person. A verdict is a probability attached to a file, produced by a stated method on audio of a stated quality. Every term here exists to keep that sentence from being rounded up into a finding.

Defined here rather than given a page of their own:

  • Detection rate. The share of synthetic clips a system correctly flags, measured only on generated audio. It says nothing about how often genuine speech is wrongly flagged, because that is computed on a different set of files.
  • False negative. A synthetic recording the system passed as human. The other error direction, and the one a single headline accuracy figure usually hides alongside the first.
  • Confidence. How strongly the evidence in one particular file supports the verdict returned for it. Distinct from a detection rate, which is a statistic about a population of files and cannot tell you what to believe about yours.
  • Unclear. A verdict, not a failure. It means the clip did not carry enough usable signal to answer — too short, too compressed, too noisy, or too many overlapping speakers. Forcing those files into one of the other two answers is where a great many false positives are manufactured.
  • Attribution, and unknown generator. Naming the system that produced a clip where its signature is recognised, and declining to name one where it is not. A wrong attribution is worse than none, because it sends somebody off to accuse a company.
  • Base rate. How common synthetic audio actually is in the material being reviewed. When it is rare, most flags will be wrong even at a low false positive rate, which is arithmetic rather than a criticism of any particular detector.

FAQ

Questions about the vocabulary

Why does the difference between these terms matter?

Because each one implies a different question, and a system built to answer one cannot answer another. Speaker verification asks who is speaking; detection asks whether a machine produced the audio. A team that confuses them buys the wrong tool and trusts the wrong result.

What is the single most confused pair on this page?

Speaker verification and synthetic speech detection. A good clone matches an enrolled voiceprint, so a verification system reports a match and is not malfunctioning. It answered its own question correctly. It was simply not asked whether the audio was generated.

Is a deepfake the same as synthetic speech?

No. Synthetic speech describes how audio was produced and stops there. Deepfake adds an accusation of deceptive intent, which is not a property of a file and cannot be measured from one. Audiobooks and screen readers are synthetic speech and are not deepfakes.

Which term should I use in a report or a policy?

The neutral one. Write that a recording shows evidence of machine generation, at a stated confidence, on audio of a stated quality. Anything stronger asserts intent, identity or authorship that the analysis did not establish and cannot support under challenge.


Where the terms are put to work

A glossary is only useful if the pages it feeds use the same words carefully. The methodology sets out what the analysis measures and in what order; accuracy reports both error directions by recording condition rather than a single figure; limitations lists the conditions under which we already know a result is unreliable. The learning centre covers the same ground for people who have a specific recording in front of them and a decision to make about it.

Truthring is pre-launch. Quantitative results are forthcoming, and where a figure would appear on these pages you will find a marker rather than a number we cannot yet support.

Reviewed