Truthring
Coverage · research systems

Detecting VALL-E and its descendants

VALL-E is a research speech synthesis system described in published work by researchers at Microsoft. There is no shop for it and, as far as we are aware, no consumer product you can sign up for [VERIFY: verify]. Coverage on this page is therefore a slightly different sort of promise from the one we make about a commercial vendor, and it is worth explaining before you read a rate.

Family-level coverage Signature held since [VERIFY: date] · last retested [VERIFY: date]
Published research samples, clean[VERIFY]%
Independent reimplementations, clean[VERIFY]%
Compressed or phone-carried audio[VERIFY]%
Attributed to this family by name[VERIFY]%
Real human speech wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips drawn from [VERIFY: n] independent implementations, held out of training. What goes into the set: research dataset.


Research is not a category of harm. It is a schedule.

People sometimes read coverage of a research system as scaremongering — a detector claiming credit for guarding against something that does not exist in the world. The objection would be fair if the pipeline from a paper to a phone call were long or uncertain. It is neither.

What actually happens is dull and repeatable. A capability is demonstrated and written up. Other developers, working only from the description, produce their own versions and publish the code. Those versions get weights, then a command-line wrapper, then a graphical one, then a hosted demo that requires no installation at all. At the end of that chain sits someone who has never read a paper, using a capability the paper introduced.

Each step is public and each takes a knowable amount of time. A detection service that begins work at the last step is arriving after the harm has already reached people. We would rather hold a weak, honestly labelled signature early than a strong one that arrives late.

1

Capability is published

A method is described with enough detail to be reproduced, usually accompanied by demonstration audio. Nothing at this point is usable by a non-specialist.

Typically usable by: researchers

2

Independent reimplementation

Other developers build their own version from the description. Results vary in quality and, importantly for us, vary in the fine detail that attribution depends on. Several implementations of the same idea are not the same system.

Typically usable by: developers

3

Weights and wrappers circulate

Trained weights are shared, then packaged with an interface. This is the step that changes the population of users, and it is the step at which our attribution rates start to matter to anyone outside a lab.

Typically usable by: anyone with a laptop

4

It stops having a name

By the time the capability is embedded in an app or a hosted demo, the user has no idea which lineage they are using. The clip that arrives in someone’s inbox carries no label, which is the whole reason attribution has to be done from the audio.

Typically usable by: everyone


What a family signature can and cannot say

For a commercial vendor, a signature tracks a product: one publisher, one release schedule, one thing to retest when it changes. For a research lineage, the object is looser. Several teams implement the same described method and arrive at systems that share broad architectural traits while differing in specifics.

So our verdict here is deliberately coarser. A clip may be reported as consistent with this family, which is a genuine and useful statement, without being reported as the output of any particular build — because the traits we can read are shared by implementations we cannot separate. We would rather publish the coarser claim than dress it up.

That also means the two rates above should be read differently. The synthetic-or-not figure describes something stable: the difference between manufactured and captured audio does not depend on whose implementation produced it. The attribution figure describes something fragile, and it is the number most likely to fall when a new reimplementation appears.

Where this fails

  • Implementations that diverge sharply. A version that departs substantially from the described method may not match the family signature at all, and will be reported as unknown generator if it is reported as synthetic.
  • Demonstration audio that has been re-encoded. Research samples circulate as compressed files copied between sites. By the time one reaches us it may have been through several encodes, each removing evidence.
  • Small sample sizes. Coverage of research families rests on fewer clips from fewer sources than coverage of a shipped product. Confidence intervals are correspondingly wider; the figures are on the benchmark page.
  • Short clips. Below [VERIFY: n] seconds the family signature contributes little and the verdict rests on capture history.

The asymmetry worth understanding. Likely synthetic is the stronger verdict, because reaching it requires positive evidence in the file. Likely human is weaker: it may mean the recording is genuine, or that compression and re-encoding destroyed the evidence on the way. For research-lineage audio, which usually reaches us second- or third-hand, the weaker reading is the more common one.


How this shapes what we choose to cover

We add a system to our coverage when we can measure it, not when it becomes newsworthy. In practice that means we monitor published work, build test material as reimplementations appear, and publish a page like this one with the honest label — family-level rather than product-level — rather than waiting for a launch that may never come.

It also means we sometimes carry a signature that produces very few real-world matches. That is an acceptable cost. The alternative is a coverage list that describes the market rather than the threat, which is a comfortable thing to publish and not much use to anyone holding a suspicious file.


Questions

Can I use VALL-E myself?

Not as a product you can buy [VERIFY: verify]. What circulates publicly are independent reimplementations built from the published description, and they are what our measurements are taken on.

If it is not a product, why does it appear in your coverage list?

Because coverage tracks capabilities that reach people, and research reaches people by a well-worn route. Waiting for a commercial release would mean arriving after the audio is already circulating.

Will a verdict name VALL-E?

It will report that a clip is consistent with this family when the evidence supports that. It will not claim a specific implementation, because independent implementations of the same method are not separable by the traits we read.

Are your rates here as reliable as for commercial systems?

The synthetic-or-not rate is comparably reliable. The attribution rate is not, and it is measured on a smaller and less representative sample. Both are published rather than blended into one flattering number.

Does covering research encourage misuse?

The capability is public whether or not we describe it. What is not otherwise public is a measured statement of how well it can be recognised afterwards, which is the part that helps the person on the receiving end.

Signature last retested [VERIFY: date] against [VERIFY: n] independent implementations. Family-level coverage; attribution to a specific build is not offered.

Reviewed