Detection as licence enforcement
Replica Studios is a commercial speech company working in games and performance production. The checks people run on its output are unlike almost everything else on this site: there is usually no victim, no scam and no disputed identity. There is a performer, a studio, an agreement between them about what may be generated, and a question about whether the shipped audio stayed inside it.
Signature held · last retested [VERIFY: date]| Audio you are checking | Reads as synthetic | Named as Replica |
|---|---|---|
| Delivered line, source quality | [VERIFY]% | [VERIFY]% |
| Shipped asset, compressed for a memory budget | [VERIFY]% | [VERIFY]% |
| Cinematic dialogue under a music and effects mix | [VERIFY]% | [VERIFY]% |
| Captured from gameplay footage | [VERIFY]% | [VERIFY]% |
| Human performance wrongly flagged | [VERIFY]% | — |
Measured on [VERIFY: n] lines across [VERIFY: n] voices and [VERIFY: n] delivery formats, generated on [VERIFY: date] and held out of training. Composition: benchmark.
The question is consent, and the audio does not contain it
Voice work in games has always involved a performer signing away specific uses of a specific recording. What changed is that a session can now produce not only lines but the capacity to generate lines — including lines nobody performed, in a sequel nobody has announced, years after the studio and the actor stopped speaking.
So the agreements got longer, and the questions in them got sharper. May the session material be used to build a model? For this title only, or for the franchise? Does approval attach to each new script, or was it given once? What happens on a re-release, a localisation, a piece of downloadable content? Performer representatives have spent a great deal of effort on those clauses.
Here is the awkward part. Every one of those questions is about permission, and permission leaves no trace in a waveform. A line generated entirely within the terms and a line generated in breach of them are, to the analysis, the same object. There is no acoustic difference between licensed and unlicensed, because the difference exists in a document.
What detection supplies is the fact that makes the document matter: this audio was generated. Without that, a performer raising a concern is describing an impression — that sounds like me and I do not remember recording it
— which is easy to wave away. With it, the conversation moves to the terms, which is where it can actually be resolved.
How the check fits into an enforcement process
Preserve before you argue
Capture the disputed audio at the best quality available and record where it came from and when. A gameplay capture is worth less than a shipped asset, which is worth less than a delivered file. Note which one you have; it determines how much weight the reading can carry.
Establish production, not identity
Get a reading on whether the audio was generated. Resist the temptation to also ask whether it is your voice — that is a different discipline and a claim we do not make. Keep the two apart or the whole submission becomes contestable.
Put it against the terms
The reading only becomes enforcement when it is read next to the clause it may breach. Whether the model was permitted, for which titles, and with what approval step, is the substance. The audio evidence is the reason anyone opens the file.
Expect the process to be routine
The healthiest version of this is dull. Studios that use synthesis under licence increasingly want their own spot checks, so that a shipped build can be shown to contain only what was cleared. Enforcement that runs before release is worth more to everybody than enforcement that runs after.
Why game audio is hard on us
Shipped dialogue is among the most heavily processed speech anyone submits. Games carry thousands of lines under a memory budget, so audio is compressed aggressively and often at a lower sample rate than a podcast would tolerate. It is then mixed under score and effects, and in many titles it is processed further at runtime — a radio filter, a helmet, reverb matched to the room the player is standing in.
All of that removes exactly the fine structure the analysis reads, and some of it actively imitates the artefacts we look for. A voice put through an in-engine radio effect has had its recording chain rewritten by the engine. We would rather tell you that plainly than publish one flattering number measured on source files and let you apply it to a capture from a livestream.
The consequence for anyone building a case: ask for the delivered audio. A studio in a licensing dispute normally has it, and a request for it is reasonable. The difference between a source file and a gameplay capture is the difference between a reading you can rely on and a reading that mostly reflects the codec.
The asymmetry, applied to a licence dispute. Likely synthetic is the stronger verdict — it is reached only on positive evidence, and it is the finding that makes a contractual question live. Likely human is weaker, and on shipped game audio it is very weak indeed, because the pipeline destroys evidence as a matter of routine. A studio should not treat a human reading on a compressed asset as a clearance.
Questions
Can you prove a game used my voice without permission?
We can support the part of that claim about production — whether the audio reads as generated, and with what confidence. Permission lives in the agreement, and identity is a separate kind of analysis we do not offer. Those three things get confused constantly, and separating them makes a claim stronger, not weaker.
Does Truthring say whose voice was cloned?
No. We report how audio was produced. Matching a synthetic voice to a particular speaker is speaker-similarity work, with its own error characteristics and its own expert witnesses.
Our studio licences synthesis properly. Can we self-audit?
Yes, and this is the use we would most encourage. Checking a build against the set of voices you cleared, before it ships, catches the ordinary failure — a placeholder line that was never replaced, an asset pulled from an older project — while it is still cheap to fix.
Why did the same line score differently in two builds?
Different compression settings, a different mix, or runtime processing applied in one context and not another. Compare files at the same stage of the pipeline, and record which stage each came from.
Is a detection result enough to send a legal letter?
On its own, no. It is a probability with a stated method and a stated error rate, which makes it one exhibit among others. It is enough to justify asking a studio for the delivered audio and the clearance records, and that request usually settles the matter faster than the reading does.
Signature last retested [VERIFY: date] against Replica Studios model version [VERIFY: verify]. Rates on this page are re-measured monthly and move when the vendor ships.
Reviewed