Truthring
Coverage · real-time speech

Detecting Cartesia

Cartesia is a commercial speech company whose work emphasises low latency — speech produced fast enough to hold a live conversation rather than narrate a finished script. That is a genuine engineering achievement and it is also the development that turns voice fraud from a recorded trick into a two-way conversation. This is one of the generators where our rates are weaker, and the page says so before it says anything else.

Partial coverage Signature held since [VERIFY: date] · last retested [VERIFY: date]
Rendered file, unmodified[VERIFY]%
Streamed audio captured at the endpoint[VERIFY]%
Recorded live call, full duration[VERIFY]%
Single conversational turn under [VERIFY: n] seconds[VERIFY]%
Named as Cartesia specifically[VERIFY]%
Real human speech wrongly flagged[VERIFY]%

Measured on [VERIFY: n] clips, generated on [VERIFY: date] and held out of training. The live-call figures are measured on [VERIFY: verify — how call audio was simulated or collected]. Composition and method: accuracy and benchmark.


What latency was buying, and what it costs us

A narration system can take its time. It receives a whole sentence, considers it, and renders the audio when it is ready. A conversational system cannot: it has to begin speaking before it knows how the sentence ends, produce audio in small pieces as it goes, and be prepared to stop mid-word when the person on the other end starts talking.

That architecture changes the evidence available to us in three ways at once, and the three compound rather than add.

  • The audio arrives in pieces. Speech generated incrementally is stitched from short segments produced under time pressure. Some of our strongest cues are statistical properties measured over a long, continuous stretch, and a stretch assembled from fragments behaves differently even when it sounds seamless.
  • The turns are short. Conversation is made of a few seconds at a time. Our second and third passes need material, and a four-second answer is not much material. This is why the per-turn row above is so much worse than the full-call row: it is the same speech, cut the way a conversation actually cuts it.
  • The channel is a live call. Real-time audio is carried by codecs built for latency rather than fidelity, and they discard exactly the fine structure our first two passes read. A clip that is comfortable as a file may be marginal as a call.

Why our third pass loses its grip here

The pass that examines prosody under stress exists because generated voices historically fell apart at the edges of speech — a laugh, a restart, a sentence abandoned halfway, someone talking over you. Those moments were where a system trained on clean read speech had nothing to draw on.

Conversational synthesis is trained on and optimised for exactly those moments. Handling interruption smoothly is not an accident of the design, it is the product. So the behaviour that used to be our tell has become the thing the vendor competes on, and the pass that once carried a real-time verdict now contributes less than it does on a narration system.

What remains is production evidence rather than performance evidence: whether the audio was captured through a physical chain, and whether it carries the regularities of a particular pipeline. Both survive a live call less well than we would like, which is the honest summary of this page.


What to do instead of listening harder

Because our figures here are weaker, the advice on this page leans on something other than the product. A conversation is decided while it is happening, and no analysis of a recording arrives in time to help the person having it.

  1. End the call and dial back. Use a number you already hold, not one you were just given. A system generating speech into a call cannot answer a phone you dialled. This single habit defeats nearly every real-time voice attack, regardless of how good the synthesis is.
  2. Do not treat interruption as a test. Talking over the caller to see whether they stumble was once informative. It is now close to a demonstration of the vendor’s roadmap.
  3. Do not treat urgency as a reason to skip the callback. Manufactured time pressure is the mechanism, not a side effect. See how voice scams are run.
  4. Record if you lawfully can, and keep the original. A verdict afterwards will not protect you, but it may protect the next person and it may matter to an investigation.

Where this fails

  • Short turns. The dominant limitation. Submit the longest continuous stretch of the other party speaking, not the most alarming sentence.
  • Two speakers in one file. A mixed call recording produces a blurred whole-file verdict. Where you can, submit a channel-separated recording or an excerpt of one speaker.
  • Low-latency codecs. Real-time transport discards more than a stored file does, and it discards the parts we use.
  • A model newer than our last retrain. This field is moving faster than the narration field. Attribution here decays sooner than it does elsewhere on the site, and unknown generator is a common and correct outcome.

The general list: where Truthring is wrong.

The asymmetry does more work here than anywhere. Likely synthetic is the stronger verdict, because positive evidence had to survive a hostile channel to be found at all — a synthetic finding on degraded call audio is worth taking seriously. Likely human is the weaker verdict, because it can also mean the evidence was destroyed in transit, and on real-time call audio destruction in transit is the normal case rather than the exception. Do not read a human-leaning result on a live call as reassurance.


Questions

Why is this page’s number lower than the others?

Because the material is harder and we would rather publish that than average it away. Short turns, incremental generation and low-latency transport each cost us something, and they occur together.

Can you analyse a call as it happens?

[VERIFY: verify — whether live analysis is offered]. The structural point stands regardless: a conversation is decided in real time, and a post-hoc verdict helps the next person rather than the current one.

How do I tell mid-call that I am talking to a machine?

Stop trying to tell. Hang up and dial back on a number you already have. That works against synthesis of any quality, which is more than we can say for listening carefully.

Does a long pause mean it is a machine?

No. Latency varies with network conditions for humans too, and reducing latency is the entire point of this class of system. Treat pauses as evidence of a bad connection.

Is it worth submitting a call recording at all?

Yes, with the reading rule above applied. A synthetic finding on call audio is informative. A human-leaning one mostly tells you about the channel.

Signature last retested [VERIFY: date] against output rendered on [VERIFY: date]. Rates on this page are re-measured monthly and move more than most, because this part of the field ships more often.

Reviewed