Truthring
Glossary · what detection measures

Spectral analysis

The standard way of looking at speech, and the most misused picture on the internet. A spectrogram shows a great deal about a file’s history and almost nothing you can read as proof by eye.

Definition

Spectral analysis examines how a recording distributes its energy across frequency, rather than how its amplitude changes over time. It is the usual way of looking at speech, because the structure that distinguishes one vowel, one voice or one reconstruction method from another lives in the frequency domain and is invisible in a waveform.


What you are looking at when you look at a spectrogram

A waveform is a picture of pressure against time. It shows loudness and silence and essentially nothing else; two completely different vowels at the same volume look alike. Decomposing short slices of the audio into their frequency content and stacking them produces a spectrogram: time along one axis, frequency along the other, energy as brightness.

In voiced speech the vocal folds vibrate at a fundamental rate, producing a stack of harmonics above it. The shape of the throat, mouth and nose then emphasises certain frequency regions and suppresses others, and those resonances — formants — are what make an ee an ee rather than an ah. They also move continuously, because a tongue has mass and cannot teleport between positions. Much of what distinguishes real speech from reconstructed speech is in how that structure behaves, not in whether it is present.


Why generation can leave traces here

Most current systems do not produce a waveform directly. They produce a compact intermediate representation of the sound and then reconstruct an actual waveform from it, and reconstruction is where information has to be invented.

The traces that result are of a general kind: energy that stops behaving naturally at particular band boundaries, harmonic structure that is too orderly or too smooth in its transitions, high-frequency detail that carries less of the irregularity a real vocal tract and a real room produce. None of that is a single fingerprint, and none of it is visible as a shape somebody can point at. It is a statistical argument made across a whole clip, which is why the method page lists spectral structure as one family of measurements among several rather than as the answer.


The telephone removes the part you wanted

A phone connection carries a narrow slice of the audible range, roughly the band that keeps speech intelligible, and discards the rest before the audio ever reaches you. A large share of the spectral evidence detection depends on sits in what was discarded.

This is the single most important practical fact about the discipline, and it is why we treat compressed telephone audio as the real case rather than the awkward one. A detector benchmarked on studio recordings has been measured in a band nobody’s voicemail ever passes through. The phone audio note works through what survives.


Commonly confused with: reading a spectrogram as proof

A spectrogram posted with a circle drawn on it is not evidence of synthesis, and this is the most common misuse of the term. What the picture genuinely shows is history: a sharp horizontal edge where a codec cut the top off the audio, a vertical discontinuity where two takes were joined, a region of perfectly flat silence that no microphone produces, a change in noise floor mid-clip. Those are useful, and they belong to audio forensics.

What it does not show is generation. The features that matter there are distributional and subtle, and the human visual system is extremely good at finding patterns in textured images that mean nothing. Its neighbour on the other side is prosody, which measures the same recording along the time axis instead — rhythm and intonation rather than frequency structure.


How it is weighted in a Truthring result

Spectral measurements contribute more when a recording still has bandwidth to measure, and less when it does not. A clip that has been through two rounds of compression on its way through messaging apps has lost much of what this family of features reads, and the confidence attached to the verdict is reduced accordingly rather than quietly held constant.

That is also why a likely human result deserves less weight than people give it. It can mean no generation evidence was present. It can equally mean the band that would have carried the evidence was thrown away before the file arrived. The limitations page lists the conditions under which we already know this happens.


FAQ

Questions this term raises

Can you see a deepfake in a spectrogram?

Not reliably, and pictures posted online claiming to show one are usually showing something else: a codec cut-off, a splice, or ordinary compression. Generation traces are statistical and subtle rather than shapes a person can point at.

What does a spectrogram actually tell you?

Mostly a file’s history. Where the audio was band-limited, where separate takes were joined, whether silence is digitally flat, whether the noise floor changes mid-recording. That is forensic information about handling, not proof of how the speech was produced.

Why does phone audio make spectral analysis harder?

Because a phone connection carries only a narrow band and discards the rest, and much of the evidence detection relies on sits in what was discarded. A detector measured on studio recordings has been tested in frequencies a voicemail never contains.


Reviewed