Truthring
Coverage · dubbing and translation

A voice speaking a language it never spoke

Camb.ai is a commercial speech company working in translation and dubbing. Its output is a particular kind of object: a real person’s voice, recognisably theirs, delivering words in a language they may not speak at all. Nothing about the voice is invented. Everything about the utterance is.

Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
Clean dub, well-covered language[VERIFY]%
Clean dub, thinly covered language[VERIFY]%
Dub inside a platform-encoded video[VERIFY]%
Named as Camb.ai specifically[VERIFY]%
Real speech in a second language wrongly flagged[VERIFY]%

Measured on [VERIFY: n] dubbed clips across [VERIFY: n] language pairs, generated on [VERIFY: date] and held out of training. Language coverage is listed at limitations.


Why our rates are worse in some languages, and why we say so

A detector learns what real speech looks like and what generated speech looks like, and it learns both from examples. If it has heard a great deal of one language and little of another, it will be confident in one and guessing in the other, while producing output that looks identical in both.

That imbalance is not a quirk of our dataset. It is the shape of the whole field. Speech corpora are heavily skewed towards a handful of widely resourced languages; so is the synthetic material available to train against; so is the published research everyone builds on. We inherited that skew and have not fixed it.

The result is that our figures outside [VERIFY: languages] are weaker, and weaker in a way that is worth understanding precisely. Two distinct things degrade. The recording-chain analysis holds up reasonably well — a microphone and a room behave the same regardless of what is being said. The prosody analysis does not, because it depends on knowing how speech in that language ordinarily moves, and if we have not learned the ordinary, we cannot recognise the odd.

There is a second, less obvious cost. False positives rise. A real recording of a person speaking a language they learned as an adult — hesitations in unusual places, stress patterns carried over from their first language, a careful over-articulation — can look, to a model trained mostly on fluent native speech, like something manufactured. We would rather flag this openly than let someone use a confident number in a language where it does not hold.

What dubbing changes about the question

In a same-language clone, an audience retains some ability to judge. People who know the speaker have heard them talk for years; a phrasing that is not theirs, a rhythm that is slightly off, can register even when the timbre is right. That intuition is unreliable, but it is not nothing.

A dub removes it. If you are listening to a public figure speak a language they do not speak, you have no baseline at all. You cannot know whether they would phrase a thought that way, because they have never phrased a thought in that language. Every stylistic oddity is explained by the translation. The listener’s only remaining check is the timbre — which is precisely the part that was preserved on purpose.

This is why translated clips are useful for a certain kind of misinformation. A statement attributed to a foreign official, circulating among an audience who cannot assess the original, is difficult to challenge and easy to spread. The counter is not a better ear. It is finding the source recording and checking what it actually says.


Dubbing is mostly ordinary work

Almost all of this is legitimate, and it is worth keeping the proportion straight. Localising a training course, a documentary or a product announcement into a dozen languages in the presenter’s own voice is a real improvement over a subtitle or a different voice actor per market. Audiences generally prefer it. Broadcasters use it. The presenter usually asked for it.

So a synthetic verdict on a dubbed clip is not an accusation, and the useful question is not was this generated — of course it was, that is the product — but was the audience told. Disclosure practice varies widely by market and platform, and where it exists it usually attaches to the publisher rather than to the audio. [VERIFY: verify the disclosure rules that apply in your market.]

The cases that need attention are narrower: a dub of a person who did not consent to being dubbed, a translation that changes the substance of what was said, or a clip presented as an original recording rather than as a localisation. Only the first and third leave any trace we can read.

How to read a verdict on a dub. Likely synthetic is the stronger answer, reached only on positive evidence — and on a dub it is close to expected, since the audio was generated by design. Likely human is the weaker answer, and in a thinly covered language it is weaker still: it may mean a person spoke, or it may mean we did not have the material to know better. Treat a human reading in a language outside our well-covered set as close to no information.

If you are checking a translated clip

  • Find the original-language recording first. What the speaker actually said is a stronger and cheaper check than any analysis of the dub.
  • Tell us the language when you submit. It changes which confidence figure applies, and a result reported without it is less useful to you.
  • Prefer the longest continuous stretch of speech. Dubs are often assembled per sentence and short excerpts can behave inconsistently.
  • If the clip is a public figure making a newsworthy statement, treat the provenance of the file — who posted it, where it first appeared — as more decisive than the audio. Detection is slower than a chain of custody.
  • For anything where money, safety, employment or a legal claim turns on the answer, use the result as one input and verify through a channel you already trust.

Questions

Can AI dubbing be detected?

In the languages our training covers well, yes — a share of clean held-out clips that has not been measured yet read as synthetic. Outside [VERIFY: languages] the figure is lower and the confidence we report drops with it, deliberately.

Which languages are you actually good at?

[VERIFY: list the well-covered languages.] We publish the list rather than a single global accuracy number, because a single number would flatter the languages we have least data for.

Could my genuine recording be flagged because I have an accent?

It is possible, and it is one of the failure modes we watch most carefully. Speech in a second language can carry patterns that a model trained mainly on native speech reads as unusual. If you believe this has happened, send the reference code and the recording and we will look at it directly.

Is a dubbed clip a deepfake?

Not by default. A disclosed localisation and a fabricated statement can be produced by the same tool and look the same to the analysis. The difference is what the audience was told, and that is context rather than signal.

Can you tell me whether the translation is accurate?

No. We look at how audio was produced, not at what it means. Comparing a dub against the original transcript is a translation question, and for anything consequential it needs a human translator rather than a detector.

Signature last retested [VERIFY: date] against Camb.ai model version [VERIFY: verify]. Per-language rates are re-measured monthly and differ substantially from the headline figure.

Reviewed