Truthring
Coverage · text to speech

Detecting Murf AI

Most people who ask about Murf are not investigating a fraud. They are sitting in front of a training module, an onboarding film or a quarterly update, and they have noticed the narrator never quite takes a breath. The technical question — can this be detected — has a short answer: on clean exported audio, [VERIFY]%, named specifically in [VERIFY]% of those. The question underneath it is harder and is the reason this page exists.

The question people actually type

Does an internal video have to say the voice was synthesised?

We cannot give you a legal answer and would not trust one you found on a vendor’s website. What we can offer is the distinction that most organisations end up drawing once they have argued about it for a while, because it is the distinction that maps onto what a detector can and cannot see.

Narration falls into two kinds. The first is functional: a voice exists so that written material can be heard rather than read. Nobody is being represented. The voice is closer to a subtitle than to a performance, and its synthetic origin is not information anyone is being denied. The second is attributed: the audio is presented as a particular person speaking — a named colleague, a director, a safety officer, someone whose authority is part of why you are meant to listen. Here the synthesis is doing something the first kind is not. It is making a claim about who is in the room.

Detection cannot tell those two apart. It reports how audio was produced. Whether production that way was appropriate depends on what the video claimed, and no analysis of a waveform can recover a claim made in a slide, an email or a meeting invitation.

Usually uncontroversial

  • Procedure walkthroughs and compliance modules with a generic narrator
  • Localised versions of an existing film, where the original narrator was also generic
  • Interface prompts, tutorials and product demos
  • Drafts and internal reviews before a human record is made

Usually needs saying out loud

  • Anything presented as a named individual speaking
  • Leadership messages, particularly about jobs, restructuring or conduct
  • Safety, medical or legal instruction where a person is being trusted, not just heard
  • Material a former employee’s voice appears in after they have left

Detected Signature held since [VERIFY: date] · last retested [VERIFY: date]
Exported narration track[VERIFY]%
Audio pulled from a published video file[VERIFY]%
Screen recording of a learning platform[VERIFY]%
Named as Murf specifically[VERIFY]%
Real human narration wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips across [VERIFY: n] voices, generated on [VERIFY: date] and held out of training. Composition and method: accuracy and benchmark.


Why corporate video is comparatively easy audio

Murf AI is a commercial text-to-speech company aimed at business users making narrated material. The clips we see from it arrive in better condition than almost anything else we handle, for reasons that have nothing to do with the generator and everything to do with the workflow around it.

Narration is exported once, at a sensible bitrate, into an editing timeline. It is not shouted across a room, not squeezed through a phone network, not forwarded five times before it reaches you. Its recording chain is absent in the plainest possible way, because nothing was ever recorded — the file was rendered. Our first pass therefore does most of the work here without much help from the other two.

The second pass, which is what allows a verdict to name the system rather than merely call it synthetic, benefits from something else: business narration is deliberately consistent. A voice is chosen for a series and used across dozens of modules at the same settings. Repetition is what signatures are made of.

The third pass contributes least. It looks at how a voice behaves when speech gets messy — interruptions, restarts, laughter, a sentence abandoned halfway. Scripted narration contains none of that by design, so on this material the pass has almost nothing to examine. That is a limitation worth stating plainly rather than hiding in a footnote.

Where this fails

  • Platform re-encoding. Learning management systems and internal video portals re-compress aggressively. A file that would have been comfortable at export can arrive stripped of the detail the analysis reads.
  • Screen recordings. Capturing a module by recording your own screen, or worse by pointing a phone at it, introduces a genuine recording chain and undoes the strongest pass.
  • Music and mixing. A narration bed under a soundtrack, ducked and compressed by a video editor, is harder than the same voice bare. Submit the cleanest excerpt you have, ideally one where the narrator is speaking alone.
  • Mixed narrators. Many corporate films alternate a synthesised narrator with recorded interviews. A whole-file verdict on that is meaningless. Cut the section in question.

The general list, which applies to every system rather than this one: where Truthring is wrong.

The asymmetry, and why it matters in an HR context. Likely synthetic is the stronger verdict, because it is reached only when positive evidence is present. Likely human is weaker, because it can also mean the evidence was destroyed in transit — and internal video pipelines destroy evidence routinely. A human-leaning result on a re-encoded module is not a finding that a person narrated it. It is the absence of a finding, and it should never be written into a process as though it were an exoneration or an accusation.


Questions

Does an internal video have to disclose a synthesised voice?

It depends on what the narration is being asked to carry, and on rules that vary by jurisdiction and sector — check what applies where you operate [VERIFY: verify]. The practical line most organisations settle on is between a voice that merely reads and a voice that is presented as somebody.

My voice appears in a company video. What can a detector tell me?

That the audio was generated rather than recorded, which is the factual half of the complaint. It cannot tell you whether you agreed, whether a licence exists, or whether the voice was modelled on yours at all as opposed to simply resembling it.

Can we use a report in a disciplinary process?

As one exhibit, with its stated error rate attached and its silence about intent acknowledged. Not as the case. A result carries a reference code precisely so the person it concerns can have the analysis re-run and contested.

We localise our training films into eight languages. Is that a disclosure problem?

Rarely, when the original narrator was already generic and nobody is represented as speaking. It becomes one the moment a real person’s voice is carried across languages they never spoke.

Should we just label everything?

It is the cheapest defensible position and it costs a line of text. Organisations that adopted a blanket label before an incident have generally found the argument easier than those who wrote a policy in the week after one.

Signature last retested [VERIFY: date] against Murf voice set [VERIFY: which voices]. Rates on this page are re-measured monthly and change when the vendor ships.

Reviewed