Truthring
Glossary · what the attacks are called

Vishing

A word about a channel that keeps being read as a word about a technology. That misreading sells detection tools to organisations whose actual gap is a payment procedure.

Definition

Vishing is social engineering carried out over a voice channel — a phone call or voice message that manipulates someone into transferring money, disclosing credentials or granting access. The channel defines the term, not the tooling. Most vishing involves no synthetic voice, and a synthetic voice is not required for it to work.


The word names the channel

It sits in a family of terms split by delivery: email, text message, voice. The split is useful because each channel changes the victim’s situation rather than the attacker’s script.

Voice is the strongest of the three for one reason above all others: it runs in real time. An email can be read twice, forwarded to a colleague, and hovered over to see where the link goes. A call has to be answered now, cannot be shown to anyone without ending it, and produces no artefact anybody else can examine. Every pressure technique that fraud has used for a century works better when the target cannot pause.

Tone does the rest. Authority, impatience, warmth and distress all travel down a phone line intact, and they are far harder to disbelieve than the same words on a screen.


What synthesis changed, and what it did not

The scripts are unchanged. An urgent payment, a compromised account, a supplier changing bank details, a relative in trouble, a security team asking you to read out a code. None of these needed a machine and none of them were invented recently.

What changed is that the voice can now be a particular person’s. That removes a defence people were relying on without noticing — the assumption that a familiar voice was evidence of a familiar person — and it lets an attacker start warm rather than cold. It also raises the ceiling: the same pretext can now be aimed at someone who knows the person being impersonated well.

The countermeasure did not move at all. Verification through a channel you initiated defeats both the old version and the new one, because neither an impersonator nor a clone controls the number already in your phone. That is why our practical material is built around procedure rather than perception, from the callback card to the payment verification policy.


A concrete case

A finance assistant takes a call from someone at a long-standing supplier, who mentions the invoice number correctly and explains that their bank has changed. The call is friendly and unhurried. There is no synthetic voice anywhere in it, and the details came from an invoice attached to an earlier compromised email.

This is ordinary vishing, and it is what most of it looks like. A detector would have had nothing to analyse, since there is no recording and nothing generated. The control that stops it is a rule that bank details are never changed on the strength of an inbound call, applied to everyone without exception so that no individual has to be the one who doubted a familiar contact.


Commonly confused with: deepfake audio

The phrase “AI vishing” runs the two together and quietly reframes a procedural problem as a technical one. Deepfake audio describes a recording. Vishing describes an interaction. A vishing call can use a cloned voice, a converted voice, an impersonator, or nothing but a plausible manner.

The distinction has budget consequences. An organisation that reads its exposure as a synthesis problem buys analysis tools that run on files after the event. An organisation that reads it as a channel problem changes how instructions are authorised, which stops the version with a cloned voice and the version without one at the same time.


Where analysis fits afterwards

Detection is a post-mortem instrument here. It runs on a recording, and a live call is not a recording. If a voicemail or voice note survives, running it can help establish what happened, whether the voice was generated, and whether to warn other people who might get the same call — but that is investigation, not defence.

The article that goes with this definition is our guide to voice phishing, which covers how the calls are structured and what to do during one. How voice scams work takes the wider view across pretexts.


FAQ

Questions this term raises

Does vishing require an AI voice?

No. Most of it is a person with a script and a plausible manner. Synthesis raises the ceiling by letting an attacker sound like someone specific, but the pretexts, the pressure and the payment routes are unchanged and long predate it.

How is vishing different from phishing?

Only by channel. Phishing arrives in writing and can be re-read, forwarded and checked. Vishing arrives as a live voice that has to be answered immediately, leaves no artefact to show a colleague, and carries authority and urgency in the tone.

Can a detector stop a vishing call?

No. Analysis runs on a recorded file after the fact, and a call in progress is not a file. What protects you during the call is ending it and dialling back on a number you already hold. The detector is useful afterwards, for working out what happened.


Reviewed