Detecting Resemble AI
Resemble AI is a commercial speech company that clones voices and also publishes detection tooling. That puts one organisation on both sides of the same question. It is a legitimate business position and it is worth thinking about before you decide whose verdict to rely on.
Measured on [VERIFY: n] clips from [VERIFY: n] distinct cloned voices, generated on [VERIFY: date] and held out of training. Dataset composition: benchmark.
The same company on both sides
Most of the systems on this site do one thing: they make speech. Resemble AI is one of a small number that also ship detection. The generator and the detector come from the same engineering culture, the same data pipeline, and in all likelihood the same building.
We want to be precise about what that does and does not imply, because the lazy version of this argument is an accusation and we are not making one. There is no evidence, and no reason to assume, that a vendor deliberately blinds its detector to its own output. A tool that failed on the audio its own customers produce would be worthless, and its customers would notice.
The real effect is quieter and structural. A team that builds a synthesiser has privileged access to that synthesiser — every checkpoint, every internal variant, every model that never shipped. Nobody else can train on that. So a first-party detector should be expected to be unusually strong on first-party audio and to have no particular advantage anywhere else. Its coverage is shaped by its business, not by the distribution of clips that actually turn up in a dispute.
The second effect is about incentives at the margin, not at the centre. Every detector vendor chooses a threshold: how much evidence before the verdict tips. That choice trades false alarms against misses. A company whose other product line is cloning is answering that question with a different set of commercial pressures than a company that only detects. Neither set is corrupt. They are just different, and the difference lands in the borderline cases — which is exactly where a disputed clip usually sits.
What that means when you are choosing
The practical test is not who made the detector but what question am I asking it.
- Provenance inside a pipeline you control. If you use a vendor’s synthesis to produce narration and you want to confirm which of your own files came from it, the vendor’s own tooling is often the best instrument available. It has information nobody else has.
- A clip of unknown origin, in dispute. Here the question has changed. You are no longer asking which of my files this is; you are asking whether an unknown party generated audio using any of dozens of systems, then sent it through a phone network. A detector whose depth is concentrated on one generator is a narrower instrument than the problem.
- A decision someone will contest. If the answer has to survive a challenge from the person it goes against, the first thing they will attack is whose interests the answer serves. That is a reason to have a second, unaffiliated reading — ours or anyone’s.
We are obviously not a disinterested party in this comparison. Take the reasoning, not the recommendation, and apply it to us too: our rates are published with their failure conditions precisely so you can check them rather than trust us.
Watermarks, and why they are not the end of the story
Vendors on the cloning side of this market often point to audio watermarking as the durable answer: mark the output at generation, read the mark later, question settled. Where a watermark survives, it is genuinely better evidence than statistical detection — it is a signal deliberately placed rather than an artefact inferred.
The limits are the ones you would expect. A watermark only exists on audio produced by a vendor that inserts one, and only survives the handling it was designed to survive. Re-recording a clip from a speaker, transcoding it through a messaging app, or cutting three seconds out of the middle can all remove it. Most importantly, absence of a watermark is not evidence of a human speaker — it is the default state of every human recording and also of every clip made by a vendor who marks nothing.
So watermarking and detection answer different questions. A watermark, when present, tells you who made something. Detection tells you, imperfectly, whether anything made it at all. The second question is the one a person holding a suspicious voicemail actually has.
What Truthring looks at on a Resemble clip
Recording chain
Whether the file carries evidence of a microphone and a room — a sensor noise floor, early reflections, the small handling sounds a body makes near a mic — or of having been rendered straight to disk.
Generator signature
The fine regularities that separate one synthesis family from another and let a verdict name a vendor rather than only say synthetic. This is the layer that decays when a vendor ships a new model, which is why this page carries a retest date.
Behaviour at the seams
Sentence starts, breath placement, the recovery after a stumble. Cloned voices are trained on tidy read speech and are least convincing at the untidy joins of real conversation.
Read the direction of the verdict. Likely synthetic is the stronger of our two answers because it is only reached when positive evidence is present. Likely human is weaker: it can mean the clip was captured by a person, or it can mean the evidence was destroyed in transit. A clean file that reads human tells you something. A phone recording that reads human tells you close to nothing.
Questions
Can Resemble AI voices be detected?
On clean audio, Truthring reads its output as synthetic in a share of held-out clips that has not been measured yet and names the generator in a rate not yet measured of those. Both figures fall on compressed audio, and attribution falls faster than the verdict.
Is a vendor that sells cloning and detection compromised?
Not in any sense we would assert. What is true is narrower: its detector is shaped by its own model access, so it should be expected to be strong on its own output and ordinary elsewhere. That is a coverage question, not an integrity one.
Which detector should I use, then?
For confirming provenance inside a pipeline you already run on that vendor, theirs. For a contested clip of unknown origin, at least one reading from a party with no stake in the generator. Where the answer matters, two independent readings beat one confident one.
If a clip has no watermark, was it recorded by a person?
No. That inference is the most common mistake made with watermarks. Absence is the normal state of human recordings and of any synthetic clip whose mark was stripped or never applied.
Can I ask you to re-check a clip a different tool disagreed about?
Yes, and disagreement is useful information rather than an embarrassment. Every result carries a reference code so the analysis can be reproduced and argued with. Where two tools split, the confidence figures usually explain why.
Signature last retested [VERIFY: date] against Resemble AI model version [VERIFY: verify]. Rates on this page are re-measured monthly and move when the vendor ships.
Reviewed