Detecting Deepgram Aura
Aura is Deepgram’s text-to-speech, built for voice agents — software that places or answers telephone calls and speaks to whoever is on the other end. Truthring identifies its output as synthetic in [VERIFY]% of clean test clips and in [VERIFY]% of recorded calls. Most people who arrive at this page are not investigating anything. They put the phone down thirty seconds ago and want to know what they were just talking to.
The short answer. If you still have the recording, a check can tell you whether the audio was generated, with the caveat that telephone audio is the hardest material we handle. If you do not have a recording, no product can reconstruct the call, and the useful question shifts from what was it to what did it want.
Automation is not the problem
An enormous amount of ordinary telephone work is now done by software that speaks: appointment reminders, delivery windows, prescription confirmations, outage notifications, first-line support triage, payment chasing. Nearly all of it is dull and legitimate, and much of it is better than the queue it replaced.
What has changed is that the software no longer sounds like software. A menu tree announced itself by being a menu tree; there was never a moment of wondering. An agent that speaks naturally, listens, and answers a question you did not anticipate removes that signal entirely, and it removes it for everyone at once — including people who would never have been fooled by a recorded message.
So the live issue is not synthesis. It is whether the caller told you, and whether they will tell you if you ask.
What a disclosed agent looks like
- Identifies itself as automated in the opening seconds, unprompted
- Answers “are you a real person?” directly and immediately
- Offers a route to a human without an argument
- Does not ask for credentials, payment details or one-time codes
What should worry you
- A claimed name and a personal manner, with no organisation given
- Deflection when asked directly whether it is automated
- Pressure to act now, on this call, without calling anyone back
- A request that moves money, access or identity in any direction
The disclosure question
Rules requiring an automated caller to identify itself are being introduced in a growing number of jurisdictions, unevenly and with real differences in who they bind, what counts as adequate disclosure, and what happens when a caller ignores them [VERIFY: verify — which regimes apply and to whom]. We are not going to summarise them here, because a wrong summary of a live regulatory question is worse than none.
The point worth making is a technical one about enforcement. A disclosure obligation is only meaningful if non-compliance can be shown after the fact, and the audio is usually the only artefact that survives a call. That puts detection in an unglamorous but real position: not as the thing that catches a fraud in progress, but as the thing that lets somebody demonstrate, weeks later, that a call which claimed to be a person was not one. That is the use case we think this generator will end up mattering for.
Measured on [VERIFY: n] clips, generated on [VERIFY: date] and held out of training, with call figures measured over [VERIFY: verify — codec set]. Composition and method: accuracy and benchmark.
What we get from call audio, and what the line takes away
A voicemail is our best case on this generator. It is a single speaker, uninterrupted, usually long enough to work with, and stored rather than streamed. A live two-way call is our worst: two people talking over each other on a narrowband channel, in turns of a few seconds each.
Within that, the first pass — whether the audio carries evidence of a microphone and a room — is the one that degrades least, because a synthetic agent has no room at all and a human call handler has a very distinctive one, complete with a headset, a colleague two desks away and the hum of an office. The second pass, which is what lets a report name the system rather than merely call it generated, is the one telephony hurts most, and unknown generator is a common outcome on real call recordings even when the synthetic verdict itself is firm. The third pass, prosody under conversational stress, is compromised for a reason peculiar to this category: agents are designed to survive interruption, so the edges of speech are exactly where the vendor has spent its effort.
If you want the call analysed later
- Keep the original file. Every re-encode removes evidence, and messaging apps re-encode.
- Do not play it back through a speaker and record that. It gives generated audio a genuine recording chain and undoes our strongest pass.
- Note the number displayed, the exact time, and what was asked of you. None of that is in the audio and all of it matters more than the audio does.
- If you can, submit a stretch where only the caller is speaking rather than a section where you are both talking.
- Check what recording rules apply to you before you record anything [VERIFY: verify — jurisdiction].
Where this fails
- Narrowband telephony. The dominant limitation on this page. See what a phone call does to a clip.
- Mixed-speaker recordings. A single-channel capture of a two-way call blurs both parties into one verdict.
- Very short exchanges. An agent that says two sentences before you hang up has not given us much.
- Human handlers reading a script. A person delivering a tightly scripted call in a quiet room is the false-positive case we watch most closely on this material.
The general list: where Truthring is wrong.
The asymmetry, in a consumer setting. Likely synthetic is the stronger verdict, because positive evidence had to survive the phone network to be found. Likely human is weaker, because it can also mean the evidence was destroyed in transit — and the phone network destroys evidence as a matter of course. If a check on a call comes back human-leaning, you have not learned that a person called you. You have learned that the line was lossy.
Questions
Can I just ask whether I am speaking to a machine?
Always ask. A compliant operator answers without hesitating. A deflection is not proof of anything, but it is information, and it is free.
Can a report prove the call was automated?
It gives a probability with a stated method and error rate, on the hardest audio we handle. Useful as one exhibit; not sufficient on its own.
Is every automated call a scam?
No, and treating it that way will exhaust you. The problem is automation that declines to identify itself while asking you for something.
The verdict said synthetic but could not name the system. Is that a weaker result?
It is a narrower one. The synthetic finding stands on its own; attribution is the part telephony destroys first, and on call recordings unknown generator is normal.
What if the caller was a person reading a script?
That is the case we most want to avoid getting wrong, and it is why the wrongly-flagged figure sits in the table above rather than in a footnote.
Signature last retested [VERIFY: date] against Aura output rendered on [VERIFY: date]. Rates on this page are re-measured monthly and change when the vendor ships.
Reviewed