Truthring
Coverage · music and rights

Detecting Udio vocals

Udio is a music generation service that produces tracks with sung vocals from a written prompt. Most of our coverage pages are written for someone deciding whether they are being defrauded. This one is not. The disputes that arrive with generated music are overwhelmingly about rights and attribution — who performed, who was paid, who is credited — and that changes what a detection result is for.

Singing — separate scale Signature held since [VERIFY: date] · last retested [VERIFY: date]
Isolated vocal stem[VERIFY]%
Vocal within a finished mix[VERIFY]%
Streaming-quality encode[VERIFY]%
Named as Udio specifically[VERIFY]%
Human sung vocal wrongly flagged[VERIFY]%

Sung material is measured on its own scale and does not compare with our speech figures — the reasons are set out on the Suno page. Measured on [VERIFY: n] generated tracks and [VERIFY: n] human performances held out of training.


The harm is distributed, which is why nobody acts

Voice fraud has an obvious victim and an obvious moment: money moves, or it does not. A synthetic vocal has neither. The loss is spread thinly across several parties, none of whom is decisively injured on any single track, and all of whom are waiting to see whether it is worth anyone’s time to act.

That diffusion is the whole difficulty. Each affected party has a different remedy available, a different threshold for using it, and a different view of what would count as sufficient evidence.

PartyWhat they loseWhat they can actually do
The performer imitatedControl of their voice and its associationsA publicity or personality-rights claim, where the jurisdiction provides one [VERIFY: verify for your jurisdiction]
Session and backing performersWork that was never commissionedVery little individually; collective action through a union or society
The rights holder or labelCatalogue value and market confusionPlatform claims, contractual enforcement, licensing pressure
The distributing platformExposure and moderation costPolicy enforcement on upload, and removal after the fact
ListenersA false belief about what they heardNothing, unless disclosure is required of the uploader

Note what is missing from the third column: nowhere does anyone win by establishing that a vocal was generated. Synthesis is not itself a wrong. It becomes actionable only when attached to a right — a voice, a recording, a contract, a platform rule. A detection result is therefore never the claim. It is at most one supporting fact inside a claim someone else has to construct.


How a result is actually used

In our experience the useful application is triage, not adjudication. Rights teams face far more suspect uploads than they can pursue, and the cost of pursuing one is measured in staff hours rather than in analysis fees. The value of a measured verdict is that it orders the queue defensibly.

1

Screen

Run flagged uploads and rank by confidence. This is the step where volume is handled, and where a stated error rate matters more than a high headline figure.

Question answered: which of these deserve attention

2

Preserve

Capture the file, the listing, the upload date and the account, and keep them with the reference code. Uploads disappear during disputes, and a result attached to a file nobody can produce afterwards is worth little.

Question answered: what existed, and when

3

Build the claim on rights, not on audio

The argument is about a voice, a recording or a contract. The verdict is an exhibit inside it — evidence that the performance was synthesised, offered with its method and its error rate rather than as a conclusion.

Question answered: what was infringed

4

Expect the result to be challenged

The other side will ask how the figure was produced, what the false-positive rate is on human vocals, and whether the analysis is reproducible. Every result we issue carries a reference code precisely so those questions have answers.

Question answered: does the evidence hold up


What a verdict is evidence of

That the vocal carries the production traces of synthesis, at a stated confidence, by a published method, on a named date, reproducible from the same file.

What it says nothing about

Whether a licence existed. Who granted it. Whose voice was modelled. Who uploaded the track. Whether the uploader believed they were permitted. None of that is in the audio, and no detector can put it there.

The second card is the one worth dwelling on, because the most common misuse of a result in this field is treating synthetic as a synonym for unauthorised. Plenty of synthetic vocals are entirely licensed — produced with agreement, sometimes by the artist themselves. A verdict cannot distinguish a permitted synthetic performance from an unpermitted one, and anyone building a process that assumes it can will eventually accuse a licensee.

Where this fails

  • Hybrid tracks. A human lead with generated harmonies, or a generated verse inside a recorded performance, is the hardest case in music and frequently returns an unclear verdict.
  • Heavily processed vocals. Tuning, layering and effects remove evidence in both directions and widen the uncertainty.
  • Stream captures. Multiple lossy encodes leave much less to read than a purchased download or a stem.
  • Short excerpts. A hook is not a sample. Submit the longest exposed vocal passage available.
  • Newer model releases. Attribution lags a service update by up to our retraining interval, even where the synthetic verdict is unaffected.

The asymmetry worth understanding. Likely synthetic is the stronger verdict, because it requires positive evidence to be present in the file. Likely human is weaker — it can mean a person performed the vocal, or that mastering and encoding destroyed the evidence before it reached us. In a rights dispute this matters directly: a likely human result is not a defence, and should not be offered as one.


Questions

Can you tell me whether the artist authorised this?

No. Authorisation is a fact about an agreement between people. Audio analysis has no access to it, and a service claiming otherwise is describing something it cannot do.

Is a synthetic vocal illegal?

Not as such, and not uniformly anywhere. What is actionable depends on the right engaged and the jurisdiction, which is why the useful question is which right was infringed rather than whether a machine was involved.

Why is the false-positive rate on human vocals published so prominently?

Because in a rights context that is the error that damages someone. Missing a generated track costs a claim; wrongly flagging a real performance costs a performer their credit.

Can this be run across a catalogue?

Screening at volume is what the API exists for. The caution is the same one as for a single track: use the result to order a queue, not to conclude a dispute.

Will the report name Udio?

When the vocal matches the signature we hold, yes. Otherwise the report reads unknown generator and the synthetic-or-not verdict stands on its own. We do not offer a nearest guess, and in a proceeding a guess would be the first thing dismantled.

Signature last retested [VERIFY: date] against Udio output generated on [VERIFY: date]. Sung-vocal rates are measured on a separate scale from speech and are re-measured monthly.

Reviewed