We are the wrong tool for your contact centre, and the right one for the morning after
If you want a synthetic voice caught while the caller is still on the line, buy an inline screening product from a vendor who builds for that. We work on recordings after the event — which is where the disputes, the reimbursement decisions and the control failures actually live.
Say the uncomfortable thing first
The obvious pitch to a bank is real-time detection on the inbound leg: flag the synthetic caller before the agent completes the authorisation. It is the right idea, and it is not our product. Doing it well means sitting inline in the telephony path, holding sub-second latency, handling the codec the carrier hands you, integrating with the agent desktop and the case management system, and surviving the operational load of a full contact centre. That is a different engineering problem with different vendors, and several of them are good at it. [VERIFY: name the category honestly and link your comparison page if you have one]
We take a file and return an analysis in seconds. That is useless while a call is live. It is genuinely useful once the call is a recording in your case file — and every disputed authorisation call in your estate is already exactly that.
Your exposure, stated plainly
Two distinct problems get bundled together under “voice fraud” and they need different responses.
Inbound impersonation. Someone calls your contact centre as a customer. Cloned voice, harvested personal data, a plausible story. Your controls are knowledge-based verification, device and behavioural signals, and whatever voice biometric you run. If you have voice biometrics on the inbound leg, a cloned voice is now an attack directly against a control you already depend on, and that is a control-testing problem before it is a detection problem.
Outbound impersonation of you. Someone calls your customer claiming to be your fraud team, and the customer moves money themselves. This is where the money and the reimbursement liability tend to sit in practice, and no detector you deploy internally sits anywhere near that call. What helps there is the callback discipline you teach customers, not a model.
Where we fit
Dispute and reimbursement
A customer says the authorisation call was not them. You hold the recording. An analysis gives the case handler a documented technical view rather than a judgement made by ear, and creates a dated record of how the question was answered.
Post-incident review
After a confirmed loss, run the call. Knowing whether the attacker used synthesis, and which family of system, changes what you fix — agent script, authentication factor, or nothing at all.
Fraud intelligence
Across a set of calls, attribution to the same generator family is a linking signal. It will not name an attacker, but it can tell you that eleven unrelated cases are one operation.
Control assurance
Test your own voice biometric and agent verification against synthetic audio deliberately, before an attacker does it for you, and hold the results for the second line.
The API takes a file and returns a verdict, a confidence, an attributed generator family where we recognise one, and the audio conditions that capped the result. Batch it against a case queue or call it from the fraud case tool. API documentation.
What this cannot do for you
It cannot stop a call. No inline decision, no real-time score, no agent prompt. If that is the requirement, we are the wrong vendor and we would rather you knew now.
It cannot identify the caller. We do not perform speaker comparison. Whether the voice matches your customer’s enrolled voiceprint is a question for your biometric vendor, and it has a different failure profile.
It cannot justify refusing a claim. A customer who was genuinely defrauded has no way to disprove a synthetic verdict on a recording of themselves, and contact-centre audio is the hardest condition we operate in — narrowband, compressed, often noisy. A refusal decision is made by a person who has weighed the transaction pattern, the device evidence and the customer’s account, and who can explain the decision to a regulator or an ombudsman without the phrase “the model said so”.
It cannot clear a call either. A likely human verdict on narrowband audio is weak evidence. Treat it as the absence of a finding, not as verification.
It will return unclear more often on your audio than on almost anyone else’s. That is honest behaviour on 8 kHz call recordings, and a vendor whose confidence does not drop on that material is not measuring what they claim.
The control that outperforms every detector
For any instruction that moves money, the callback on a number your side already holds beats a probability. It is faster than an analysis, it does not degrade with codec quality, and it works against a synthesis system nobody has trained against yet. Detection is for the recordings you are reviewing afterwards. The callback is for the money.
Obligations around fraud reimbursement, customer authentication and the use of automated processing in decisions affecting customers vary by jurisdiction and are moving quickly. [VERIFY: verify the current position in each market you operate in] Nothing here is legal or regulatory advice.
Questions from fraud teams
Can we run this over our whole call archive?
Technically yes via the API. Consider first whether you have a lawful basis and a proportionate purpose for analysing every customer’s voice, as opposed to the calls actually in dispute. Bulk screening of customers who have not complained is a materially different processing activity from investigating a case.
What do you retain?
Audio will be deleted after analysis [VERIFY: retention not yet set]; a one-way hash is kept so a report can be tied to the exact file. Data residency, sub-processors and DPA terms are on the security page and the compliance page. Get those reviewed by your third-party risk function before any production integration.
Do you attribute to a specific generator?
To a system family where the signature matches one we hold. Where it does not, the report reads unknown generator rather than guessing. Coverage is listed on the generator coverage page with the date each was added.
Can we tell customers we screen for AI voices?
Be careful with the wording. Saying you analyse disputed recordings is accurate. Implying that calls are screened for synthetic voices in real time would not be, and a customer who relied on that impression before losing money has a complaint worth taking seriously.
Start with the disputed calls you already hold
Run a handful of closed cases where you already know the answer. It is the fastest way to see what the tool is worth on your audio rather than on ours.
Reviewed