Detecting Rime
Rime is a commercial speech company whose synthesis aims at conversational rather than broadcast delivery — speech that sounds like a person talking, not a narrator reading. That is a modest-sounding difference and it is the one that costs us most. Truthring identifies its output as synthetic in [VERIFY]% of clean test clips, and the reason that figure is not higher is written into the design goal of the system itself.
Reading and talking are not the same act
Nearly all speech synthesis was, until recently, an imitation of reading aloud. Sentences arrived complete and were delivered complete. Pace was even. Emphasis landed where a professional narrator would place it. Nothing was ever abandoned halfway, because a script does not abandon anything.
Actual talking is not like that, and the differences are not decorative. People start a sentence, discard it and start again. They hold a vowel while deciding what comes next. They run three clauses together and then leave the fourth unfinished because the point has already landed. They breathe audibly, swallow, click, drop consonants, and put emphasis in places that make sense only given what they were thinking about a second earlier. Register drifts within a single utterance. Volume falls at the end of a thought.
A system that models those behaviours is not adding realism as a garnish. It is targeting the specific set of properties that a detector like ours had been quietly relying on.
Which pass this attacks, and what it does to the weighting
Truthring runs three passes on every submission, and only one of them is about performance.
Recording chain — unaffected
Whether the audio carries the trace of a physical microphone in a physical room. Sounding relaxed does nothing to change that a file was rendered rather than captured. We expect the first pass to carry proportionally more of the load here than on most generators. The weighting below is a design intention, not a measured result. [VERIFY: pass weighting not measured]
Generator signature — unaffected, but perishable
Fine regularities left by a particular pipeline. This is what allows a verdict to name a system rather than merely call a clip synthetic. It is untouched by naturalness, but it decays whenever the vendor ships, which is the ordinary risk on every page here.
Prosody under stress — degraded, by design
The pass that examines how a voice holds together at the untidy edges of speech. It was built on the observation that systems trained on clean read material had nothing to draw on when speech stopped being clean. A system trained on unclean material has plenty to draw on, and this pass loses most of its purchase.
The verdict does not collapse when a pass weakens; the weighting shifts and the confidence figure moves with it. That is what a lower headline number on this page represents. It is not that we cannot see these clips, it is that we are seeing them with two instruments instead of three.
Measured on [VERIFY: n] clips across [VERIFY: n] voices, generated on [VERIFY: date] and held out of training. The hesitation row is measured on a subset selected by [VERIFY: verify — selection rule]. Composition and method: accuracy and benchmark.
The heuristic this breaks for listeners
There is advice in circulation, some of it ours from earlier, that amounts to: listen for the seams. Ask a question the caller cannot have prepared for. See whether they hesitate naturally. Notice whether they breathe.
Naturalistic synthesis retires that advice. Hesitation is a modelled behaviour now, not a symptom of being alive. Breath can be placed. A stumble can be produced on purpose and produced well. Anyone still deciding what is real by how comfortable a voice sounds is using a test that the field has spent several years explicitly optimising against, and the more convincing the stumble the more likely it was authored.
The tests that still work do not involve listening at all. They involve the channel: end the call and dial a number you already hold, ask something only the real person would know and that is not recoverable from anything published, or move to a medium the other party did not choose. See how voice scams are run for the versions of this that hold up.
Where this fails
- Short casual clips. The worst combination on this page. Conversational speech carries less usable structure per second than read speech, so a five-second fragment is thinner than five seconds of narration would be.
- Voice notes. Casual synthesis is most likely to be encountered in exactly the medium that compresses hardest, and each forward re-encodes it again.
- Re-recording through a speaker. Playing a clip aloud and capturing it gives generated audio a genuine recording chain — and on this generator the first pass is carrying more of the verdict than usual, so the damage is proportionally worse here.
- Genuinely disfluent human speech. Our false-positive risk on this page sits with real speakers who hesitate a great deal: tired, distressed, unwell, or speaking a language they are less fluent in.
The general list: where Truthring is wrong. For the related but separate problem of systems built for live back-and-forth, see Cartesia.
The asymmetry, sharpened by a weaker third pass. Likely synthetic remains the stronger verdict: reaching it requires positive evidence, and on this generator that evidence has had to come from production rather than performance. Likely human is weaker still than usual here, because it can mean the evidence was destroyed in transit or that the system simply performed well enough that one of our three instruments had nothing to report. Treat a human-leaning result on casual audio as an absence of finding, not a finding of absence.
Questions
Do filler words mean a real person?
No. They are a modelled behaviour and have been for a while. A convincing stumble is now weak evidence in the wrong direction.
Is this the same as the real-time problem?
Related, not identical. Latency-driven systems cost us short turns and lossy transport. Naturalistic systems cost us the prosody signal itself, whatever the channel.
What still works when that signal goes?
Production evidence: the recording chain and the pipeline signature. Neither cares how the voice performs, only how the file came into existence.
How should I submit a conversational clip?
Longer, and unedited. Give us a continuous stretch rather than the most striking sentence — on casual speech, duration buys more than drama does.
Will you get this pass back?
Partly, and never permanently. Each retraining run recovers some ground on the behaviours a new model added, and each vendor release takes some of it away. We publish the retest date so you can see how old our footing is.
Signature last retested [VERIFY: date] against Rime output rendered on [VERIFY: date]. Rates on this page are re-measured monthly and change when the vendor ships.
Reviewed