Detecting Suno vocals
Suno is a music generation service: from a written prompt it produces a finished track, instrumentation and sung vocals together. Everything on this page concerns singing, which is a materially different detection problem from speech. The rates below are measured on sung material and reported on their own scale. They should never be quoted alongside our speech figures as though the two described the same thing.
Read this before the numbers. Sung vocals are harder to place than spoken ones, and the reasons are not marginal. If you have seen a headline detection percentage for AI voice anywhere — ours or anyone else’s — assume it was measured on speech and does not transfer to music.
Measured on [VERIFY: n] generated tracks and [VERIFY: n] human-performed recordings across [VERIFY: which] genres, held out of training. Genre composition matters here and is listed in the dataset notes.
Why singing resists the analysis that works on speech
Our speech analysis rests on a simple distinction: audio that was captured carries the history of a microphone in a room, and audio that was manufactured does not. Sung vocals complicate that distinction in a way that has nothing to do with AI, because a professionally recorded vocal has been heavily manufactured too.
A released vocal has typically been pitch-corrected, time-aligned, compressed hard, de-essed, layered with doubles, and placed in an artificial space by reverb that was never in the room. Those processes deliberately remove the small inconsistencies that our first pass reads as evidence of capture. A human singer’s finished vocal is, by the standards of the analysis, already partly synthetic-looking.
Singing also constrains the signal itself. Pitch is quantised to a scale rather than drifting freely, vowels are sustained far longer than in speech, vibrato imposes its own regularity, and phrasing is dictated by a bar rather than by breath and thought. Much of the natural variability that distinguishes a person from a model in speech is simply not present in a sung line, in either the human or the generated case.
The consequence is not that detection is impossible. It is that the same method produces weaker separation, and the false-positive risk — flagging a real singer — is higher than anything we would accept on speech. That is why the human-vocal error rate appears in the table above rather than in a footnote.
What the production chain removes
Mixing
The vocal is placed under instrumentation that overlaps it in frequency and masks fine detail. Separating a stem afterwards introduces artefacts of the separation process itself, which have to be modelled or they contaminate the result.
Mastering
Loudness processing compresses dynamic range across the whole track. Quiet detail — breath, room tail, noise floor — is raised or removed, and it is exactly that quiet detail the capture analysis depends on.
Distribution encoding
Streaming platforms encode aggressively. By the time most people can obtain a track, it has been through at least one lossy stage, and often a re-record from a video platform on top of that.
The case people actually bring us
It is almost always the same one. A track appears — on a streaming service, a video platform, a fan account — performed in the recognisable voice of a named artist who did not sing it. Sometimes it is presented as a leak, sometimes as a tribute, sometimes as nothing at all. It accumulates plays before anyone with standing hears about it.
What a verdict contributes to that situation is narrow and worth stating exactly. We can report whether the vocal carries traces of synthesis, with a confidence figure and a repeatable method. We can sometimes name the generating system.
We cannot tell you whose voice was modelled. Our analysis addresses how audio was produced, not whom it resembles; resemblance is what a listener already knows and is not the disputed question. We cannot tell you who uploaded the track, whether a licence existed, or whether the vocal was a clone at all rather than an impersonator with a good ear. A convincing human imitation and a synthetic clone are different problems, and only one of them is ours.
Where this fails
- Heavy vocal effects. A vocal drenched in processing — hard tuning, vocoding, extreme layering — can suppress the evidence on either side, in both directions. This is the biggest source of unclear verdicts in music.
- Short hooks. A few seconds of refrain is not enough. Submit the longest continuous vocal passage available.
- Stem separation on a dense mix. The artefacts introduced by separation resemble, at some scales, the artefacts we look for. We account for this and it still costs accuracy.
- Genres with heavily stylised delivery. Where human performance is already highly processed by convention, separation between the classes narrows further.
- Re-recorded audio. A track captured from a speaker gains a genuine capture history and defeats the first pass entirely.
The asymmetry worth understanding. Likely synthetic is the stronger verdict, because reaching it requires positive evidence in the audio. Likely human is weaker: it may mean a person sang it, or it may mean the production chain removed everything that would have shown otherwise. On a mastered, streamed track, a likely human result is close to no information, and we would rather you treated it that way than as a clearance.
Getting a usable result from a track
- Submit an isolated vocal stem if one exists. Nothing else improves the result as much.
- Failing that, submit the highest-quality file you hold — a purchased download beats a stream capture, which beats a screen recording.
- Choose a passage where the vocal is exposed rather than the loudest chorus.
- Note where the track appeared and when, and keep that record with the reference code. Provenance often decides these cases; the audio supports it.
Questions
Is your accuracy on songs the same as on speech?
No, and the two should never be quoted together. Singing figures are measured separately, are lower, and carry a higher risk of flagging a real performance.
Can you tell me whether it is really that artist singing?
No. We report whether a vocal shows traces of synthesis. Establishing whose voice it is, or is meant to be, is a different question and not one we answer from audio.
What if a person imitated the artist convincingly?
Then the vocal should read as human, correctly, because it is. Impersonation by a skilled singer is a real phenomenon and is outside what a synthesis detector addresses.
Does an instrumental matter?
Only as interference. Our analysis concerns the vocal; the backing is something to be worked around, and a dense arrangement lowers the rate.
Can a result get a track taken down?
Not on its own. Platforms act on rights claims and their own policies, and a detection result is one exhibit supporting such a claim rather than a substitute for making it.
Signature last retested [VERIFY: date] against Suno output generated on [VERIFY: date]. Singing rates are measured on a separate scale from speech and are re-measured monthly.
Reviewed