Truthring
Coverage · community model libraries

Detecting Fish Audio

Fish Audio is a speech synthesis platform organised around a library of voice models published by its users. That structural detail matters more than any property of the audio. When the voices are supplied by the community rather than by the operator, the question “who agreed to this voice being available” has no single place where it could be answered — and in practice it is not answered anywhere.

Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
Clean render, studio-sourced community model[VERIFY]%
Clean render, model built from low-quality source[VERIFY]%
Voice note or messaging app[VERIFY]%
Named as this platform specifically[VERIFY]%
Real human speech wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips generated from [VERIFY: n] distinct community models of varying source quality, held out of training. See accuracy.


A library is a different object from a tool

A cloning tool asks you to supply a sample. That is a private act with an obvious author: whoever uploaded the audio is the person answerable for it. A library inverts the arrangement. One person supplies the sample, publishes the resulting model, and then anyone at all can generate speech in that voice without ever touching the source recording or knowing where it came from.

The distinction shows up in three places. Reach: a model uploaded once can be used by thousands of people who would never have gone to the trouble of building it. Deniability: a user generating from a published model has done nothing but pick an entry from a list. And distance: by the time a clip causes harm, several parties stand between the audio and the person whose voice it copies, each of whom can point at another.

None of that is unique to one platform. It is a property of community model libraries generally, and it is why we treat them as a distinct category rather than as ordinary cloning products with a bigger catalogue.

Point in the chainWho could verify consentWhat typically happens
Source audio is collectedThe uploaderMaterial is taken from public video, streams, podcasts or calls
Model is trainedThe platformAn assertion of rights is made by the uploader [VERIFY: verify]
Model is published to the libraryThe platformListing proceeds; review, where it exists, is reactive [VERIFY: verify]
Anyone generates speechNobodyThe end user has no relationship to the source at all
The clip reaches a listenerThe listenerNo provenance travels with the file

The row that matters is the last one. Everything upstream can be argued about; the listener receives a file with no history attached and has to make a decision anyway.


The person being modelled is not in the transaction

Consent frameworks assume the subject is present. Here they are not. Someone whose voice appears in a community library has no account on the platform, received no notice when the model was published, and has no standing to be told when it is used. They generally discover the situation the way anyone else would — by hearing a recording of themselves saying something they did not say.

This produces a second-order problem we see often enough to name: people needing to demonstrate that a clip is not them. That is a harder ask than it sounds, and it is where the limits of any detector need stating plainly. A likely synthetic verdict supports the claim. It does not establish who generated the clip, from which model, or with what intent, and it cannot recover the identity of an uploader.

What a verdict can support

That a recording bears the production traces of synthesis, with a stated confidence, a stated method and a reference code that lets the analysis be repeated and challenged.

What it cannot establish

Who made the model, who generated the clip, whether the subject agreed, or whether any licence was in place. Those are matters of record, and the record sits with a platform rather than in the audio.


Why community models detect unevenly

A commercial vendor trains on material it controls, so its output is consistent and our rate for it is a single meaningful number. A community library is trained on whatever its uploaders had: a clean studio session in one case, a compressed livestream capture in another, a phone recording in a third.

That inconsistency reaches the output. A model built from noisy source material tends to reproduce characteristics of that noise — a room, a codec, a microphone’s colour — in everything it later generates. The effect matters because our first pass asks whether audio carries the signature of having been captured. Source noise partially imitates a capture history, which is why we publish two separate rates above rather than one average that would describe neither case.

The same property cuts the other way for attribution. Models sharing a platform’s pipeline retain enough common structure that naming the platform is often possible even when the individual model is unknown to us.

Where this fails

  • Models built from very poor source audio. The most awkward case on this page, and the one that pulls the second rate down.
  • Clips shortened to a phrase. Under [VERIFY: n] seconds the platform signature carries little weight.
  • Post-processing. Community workflows often include editing, noise addition and loudness normalisation, each of which removes evidence.
  • Newly added pipeline changes. Attribution lags a platform change by up to our retraining interval — see method changelog.

The asymmetry worth understanding. Likely synthetic is the stronger verdict because it requires positive evidence in the file. Likely human is weaker: it can mean the clip is genuine, or that the evidence was destroyed in transit — and on models trained from noisy sources, it can also mean the source noise did the destroying before the clip was ever sent.


If your voice has been modelled without your agreement

Collect the material before it moves. Save the original file and note where it appeared, when, and under what account — listings and posts are removed, sometimes quickly, and a screenshot is worth more than a recollection. Check the clip and keep the reference code; a dated result is more useful than one produced months later.

Then treat it as two separate problems. Removal is a platform matter and follows the platform’s own process. Harm — a fraud, a defamatory clip, a work you are said to have performed — is a separate track and needs its own record. A detection result is an exhibit in the second, not a solution to the first.


Questions

Does anyone check that an uploader has rights to a voice?

The uploader asserts it [VERIFY: verify]. Nothing in the chain contacts the person whose voice is being modelled, and no verification step exists that the subject could participate in.

How would I know a model of my voice existed?

Ordinarily you would not, until you heard the output. There is no notice directed at the subject, because the subject is not a party to the upload.

Can Truthring tell me which community model was used?

Usually not. We can often name the platform, because models sharing a pipeline share readable structure. Identifying an individual community model from audio alone is beyond what we will claim.

Why are two detection rates listed?

Because source quality varies enormously across community models and averaging the two would misdescribe both. The lower figure is the one to plan around.

Can a result prove a clip is not me?

It can support the claim with a stated confidence and a repeatable method. It cannot prove it, and anyone offering you proof from audio alone is overselling.

Signature last retested [VERIFY: date] against [VERIFY: n] community models spanning [VERIFY: which] source-quality bands. Rates are re-measured monthly.

Reviewed