Open-weight voice cloning
Most public discussion of synthetic voice assumes a company sits behind the audio — one that can be written to, fined, or required to check consent. For a large and growing share of the clips people actually receive, no such company exists. The model was downloaded and run on a laptop. This page is about what that removes, what it does not, and how our coverage of these systems differs from our coverage of products.
The short version. Open weights do not make audio more convincing. They remove the account, the quota, the consent prompt, the log, the billing record and the ability to withdraw access — all at once, permanently, for every copy already distributed.
What vendor-facing rules actually reach
Almost every proposed control on voice cloning is a control on an intermediary. Verify that the uploader has rights to the voice. Log who generated what. Watermark the output. Rate-limit the API. Suspend the account. Preserve records for investigators. Each of these is sensible, and each assumes there is somebody in the middle to be obliged.
That assumption holds for hosted products and fails completely for a model file on a hard drive. There is no upload to check, no request to log, no quota to enforce, no account to suspend and no record to preserve. The rule does not fail partially. It has nothing to attach to.
None of which makes vendor rules pointless. They shape the behaviour of the overwhelming majority of users, who reach for the easiest tool and stay inside it, and they make casual misuse meaningfully harder. But they select for a residue: the people willing to install software to avoid a checkbox. That residue is small and it is almost exactly the population that causes deliberate harm.
| Control | Hosted product | Model running offline |
|---|---|---|
| Consent attestation at upload | Applies, though usually unverified | No upload exists |
| Generation logs available to investigators | Applies | Nothing is recorded anywhere |
| Rate limits and cost per minute | Applies | Bounded only by the machine |
| Account suspension | Applies | No account |
| Output watermarking | Applies when implemented | Optional to the operator |
| Withdrawing the capability | Possible | Impossible once distributed |
The final row is the one that changes strategy. Every other control degrades; that one inverts. A capability released as weights cannot be recalled, and old versions keep producing clips long after anyone would call them current.
Why watermarks cannot carry the weight
Watermarking is often proposed as the answer, and it does real work in one direction. A watermark that is present is good evidence that audio was generated. The trouble is entirely in the other direction: a watermark that is absent tells you nothing at all, because it may never have been applied.
Any scheme where the generator chooses whether to mark its output is a scheme that marks compliant use and leaves deliberate misuse unmarked. If unmarked audio came to be treated as presumptively genuine, the scheme would be worse than nothing — it would hand the motivated attacker a certificate. Detection that reads the audio itself avoids this trap, because it does not depend on the cooperation of whoever produced the file.
What defenders can still rely on
Two things survive the disappearance of the vendor, and they are worth more than they sound.
The file still has a history
Audio that was captured passed through a room, a body and a microphone. Audio that was manufactured did not, unless someone went to the trouble of simulating it. That difference is in the file regardless of who ran the model, and it is what our analysis reads.
The attacker still has to reach you
A clip has to arrive through some channel, aimed at some decision. Callback on a number you already hold, dual approval on payments, and a habit of not acting on urgency all work identically against every generator, known or unknown.
The second is the more reliable of the two, and it is the one most organisations skip because it is procedural rather than technical. No detector removes the need for it. We would rather say that plainly than sell a verdict as a substitute for a process.
How coverage of these systems differs
Signatures for commercial products expire on a schedule: the vendor ships, the output shifts, we retrain. Open models break that pattern in both directions. Nothing is ever retired, so a signature stays useful long after the release it describes has been superseded. And everything branches, so new variants appear continuously as fine-tunes and forks, none of them announced.
The result is that we hold family signatures rather than version signatures for this category, and that our synthetic-or-not verdict holds up considerably better over time than our naming of a specific system does. Where a clip is synthetic but the build is unfamiliar, the report says unknown generator. We do not offer a nearest guess.
Coqui XTTS
Open-weight cloning from a short reference sample. The clearest case of a threat model with no vendor in it.
Bark
Generates laughter, breath and hesitation alongside words, retiring the heuristic most people assess recordings with.
Tortoise TTS
Slow, high-quality generation. Slowness constrains volume and does nothing about a single targeted clip.
VALL-E
A research lineage rather than a product. What reaches people are independent reimplementations of the published method.
Fish Audio
A community library of user-published voice models, many of real people, with no verification step the subject could take part in.
Everything else
The full coverage list, including the commercial systems, with per-system rates and retest dates.
The asymmetry worth understanding. Likely synthetic is the stronger verdict because it is only reached when positive evidence is present in the file. Likely human is weaker: it can mean the recording is genuine, or that the evidence was destroyed on the way — by a codec, by re-recording, or by deliberate processing. Someone running a model on their own machine controls every step after generation, so the weaker reading deserves more weight in this category than in any other.
Where this leaves policy
The useful question is not whether to regulate vendors but what to expect it to accomplish. Rules on hosted services reduce casual misuse, create records where records can exist, and give investigators somewhere to start. They will not prevent a determined impersonation, and a policy sold on that promise will be judged a failure by an outcome it was never able to affect.
What scales against the residue is downstream: verification procedures at the point where audio is acted on, and measurement of the artefact itself. Both work without knowing which model was used, or whether the person who used it agreed to anything.
Questions
If the model runs offline, can anything be traced?
Rarely from the generation itself. What can sometimes be traced is delivery — the account that sent the file, the number that placed the call, the platform it was posted to. Those records exist at the edges even when the middle is empty.
Should open release of speech models be restricted?
That is a policy question rather than a detection one, and we have no standing to settle it. What we can report is the operational consequence: once weights are distributed the capability cannot be withdrawn, so defences that assume withdrawal is possible should not be relied on.
Do you detect open models as well as commercial ones?
For the synthetic-or-not verdict, broadly comparably. For naming a system, less well, and the gap is published rather than averaged away. Per-system figures are on each page above.
Does an unwatermarked clip mean it is real?
No, and this is the most damaging misreading in the field. Absence of a watermark is absence of information.
What is the single most effective thing an organisation can do?
Remove voice from the list of things that authorise an action. If no payment, credential reset or access grant can be triggered by a voice alone, the entire category of attack loses its payload.
Signature last retested [VERIFY: date] across [VERIFY: n] open-weight families. Family-level coverage; per-system rates and retest dates are on the individual pages.
Reviewed