The candidate you interviewed may not be the person who turns up
Interview proxying is not new. What is new is that the proxy no longer needs to sound like the candidate — conversion software will do that in real time, on a laptop. A recorded-audio detector helps with part of this problem and is useless against the rest, and the difference matters before you build a process around it.
What the fraud looks like now
The recognisable version is the technical proxy: one person does the interviews, another does the job, or never does the job at all. Remote onboarding made it cheaper, and distributed teams made it harder to notice for months.
Voice conversion changed the mechanics rather than the motive. A proxy can now speak in something close to the candidate’s voice while the candidate’s video is a still frame, a poor connection or a filtered feed. Conversely, a candidate can present a voice that is not theirs at all — useful where the fraud involves a fabricated identity rather than a substituted person, which is the pattern behind the fake-worker schemes your security team is probably already briefing you on.
The signals recruiters notice first are behavioural, not acoustic: audio and video that never quite align, a candidate who declines to turn the camera on for a second interview, answers delivered with an odd latency, a reference chain that resolves to email addresses rather than switchboards, a bank detail change before the first payroll run. Those remain the strongest signals. Audio analysis is a supplement to them.
What this cannot do for you
It cannot verify identity. It has no view on whether the person speaking is the person named on the application. That is a document, right-to-work and identity-proofing question, and it needs the tools built for it.
It cannot watch a live interview. We analyse recorded files afterwards. Nothing we return can alert an interviewer during the call.
It cannot reliably catch real-time voice conversion. Conversion applied to a live human voice is a harder case than fully synthesised speech, and the video-call codec sitting between the conversion and your recording strips much of what would give it away. Treat a clean result on conference-call audio as close to no information.
It cannot tell you a candidate was dishonest. Even a strong synthetic verdict tells you something about audio, not about intent. Candidates use noise suppression, voice-isolation features and accent-clarity tools in good faith, and some of those leave traces.
It cannot reject anyone. A verdict never becomes an outcome on its own. A person reviews it, the candidate is told and given the chance to respond, and the decision rests on more than the score.
Your duties when you use it in hiring
Hiring is one of the most closely regulated places a machine judgement can appear, and the constraints are not subtle.
- Disclose before, not after. Tell candidates in the interview invitation that audio may be recorded and analysed, and why. Bury it and you convert a defensible control into a complaint.
- Human review of any negative signal. Named reviewer, recorded reasoning. A flag must not silently move an application into a rejected pile.
- No automated rejection. Decisions with significant effects on a person taken solely by automated means are restricted in several regimes, and hiring is squarely within them. [VERIFY: verify under the rules applying to each market you recruit in]
- Right of reply. If a signal contributes to a rejection, the candidate should be able to know that and respond. In practice this usually resolves the flag — the explanation is a headset or a bad line.
- Watch for disparate impact. Detection performance can vary by language, accent and recording environment, and a candidate on a poor domestic connection in a different country is not the same test case as one in a quiet office. [VERIFY: state measured performance by language, or say plainly that it is untested] A control that flags one group more often is a discrimination problem regardless of intent.
- Retention. Interview audio and any report follow your recruitment retention schedule. Do not keep a flagged recording longer because it seemed interesting.
This page touches on employment and data protection law and is not legal advice. Take advice in each jurisdiction where you hire.
Controls that work better than a detector
Live, unpredictable interaction
Ask something that cannot be pre-scripted and requires reacting to a document shared on the call. Proxies handle rehearsed answers well and improvisation badly.
Identity proofing at offer
Verify documents against a live person through a provider built for it, before contract, not after start date. This catches what audio never will.
Continuity checks
Confirm the person who interviewed is the person onboarding and the person on the first team call. Substitution usually happens at a handover, not during the interview.
Where an analysis fits is narrower than a vendor would like to admit: on a recorded screening interview where something else has already prompted concern, and on submitted audio assignments, where the candidate is speaking into their own device without a conference codec in the way. That second case is the one where the tool is actually good.
Questions from talent teams
Can we screen every interview recording?
You can, and we would advise against it. Blanket screening of all candidates is a much heavier processing activity, generates false positives at volume against people with no way to answer them, and puts a machine signal in front of hiring managers who will not read the confidence number. Use it where there is already a reason.
What do we say to a candidate who is flagged?
Say what happened, plainly: an automated analysis of the recording returned a result you want to understand, it is not an accusation, and you would like to arrange a short verification step. Most of the time the explanation is technical and the conversation ends there.
Is a poor-quality recording a red flag?
No. It is a poor-quality recording. Persistently refusing a better one, across several attempts and channels, is a behavioural signal — but that is a judgement about behaviour, not an audio finding.
Should this go in the candidate privacy notice?
Yes, and in the interview invitation as well. [VERIFY: have your employment counsel draft the wording for each market]
Use it where it is strong
Recorded assignments and follow-up screening calls, with disclosure, with human review, and never as the reason on its own.
Reviewed