Vishing
Vishing is the security industry’s term for voice phishing: a social-engineering attack delivered over a voice channel rather than by email. The word contracts voice and phishing, on the same pattern as smishing for SMS. It names the delivery channel, not the pretext, and it now covers calls that use cloned speech.
Written for security professionals, students and anyone drafting policy or training material. If you are dealing with a call right now, the consumer guide to voice phishing is the page you want instead.
Where the word comes from
The term is a compound built on an older one. Phishing borrowed the fishing metaphor and adopted the ph spelling generally traced to phone phreaking, the earlier practice of manipulating telephone signalling. When the same deception moved from email onto other channels, the pattern was reused: smishing for SMS, vishing for voice, and more recently quishing for QR codes.
Vishing entered circulation as internet telephony made outbound calling effectively free and caller identification trivially assertable. [VERIFY: earliest attested use of “vishing” with a citable source and date] The coinage was descriptive rather than technical, and that has consequences the industry still lives with, because the word names only where the attack arrives.
It is worth noting that the underlying activity is much older than any of these words. Persuading somebody to disclose a secret or authorise a payment by telephone predates the internet entirely. What the vocabulary added was a way to count and categorise incidents, not a new behaviour.
The family, and what each term actually varies
Most confusion in this vocabulary comes from mixing three independent axes: the channel an attack arrives on, who it is aimed at, and the story it tells. The -ishing family varies the first. The other common labels vary the second or the third, which is why a single incident often has several accurate names.
| Term | What varies | Delivery |
|---|---|---|
| Phishing | Channel | Email, and historically instant messaging |
| Smishing | Channel | SMS and messaging apps |
| Vishing | Channel | Telephone and voice over IP |
| Quishing | Channel | QR codes leading to a hostile page |
| Spear phishing | Targeting | Any channel, aimed at a named individual |
| Whaling | Targeting | Any channel, aimed at senior executives |
| Pretexting | Method | Any channel, defined by the invented scenario |
Read across the table and the practical point becomes visible. Vishing is not a peer of whaling; a call can be both. An attack described only as vishing tells a reader nothing about who was targeted or what was requested, and a report that records only the channel is discarding most of what an analyst would want.
What voice over IP changed
The first shift was economic. Once calls cost effectively nothing and could be placed programmatically, the telephone stopped being a low-volume, high-effort channel and became something closer to email in its economics while retaining the immediacy and social pressure of a live conversation. That combination is the reason the channel remains effective.
The second shift was structural. Caller identification was designed as a convenience feature and is asserted by the originating party rather than authenticated by the network, so a displayed number became a claim rather than an address. Every consumer-facing piece of advice about not trusting caller display exists because of this design decision. Efforts to authenticate calling numbers exist and are being deployed unevenly across jurisdictions. [VERIFY: current status and coverage of call authentication standards per market before describing them]
The third was automation. Interactive voice response systems can be imitated cheaply, so a target can be routed into what sounds like a bank’s own menu and asked to key in a card number and a PIN without ever speaking to a person. This variant is often underweighted in training material that teaches people to listen for a suspicious human.
What voice cloning changed
Before practical cloning, the attacker on a vishing call had to be a plausible stranger. That constrained the pretext space to institutional roles: the fraud department, the tax office, the courier, the support desk. Authority was the only lever available, because impersonating a specific known person was not achievable in a live conversation.
Cloning removed that constraint. A likeness good enough for a telephone channel can be derived from roughly a minute of clear recorded speech, and the material is usually already public: conference talks, podcasts, webinars, social video. The shape of that capability matters more to a defender than its mechanics. The consequence is that the pretext space now includes intimacy as well as authority, and a caller can be a chief executive, a colleague or a family member rather than an official.
Real-time voice conversion extends this further, because it lets an operator improvise, answer unexpected questions and respond to hesitation while sounding like the impersonated person. That removes the classic defensive advice of asking something a script could not anticipate.
Two defensive implications follow, and they point in opposite directions. Procedural controls survive intact: a callback to a known number, an out-of-band approval, a passphrase agreed in advance are all unaffected by how convincing the voice is. Controls that treat the voice itself as an authenticator do not survive, which is a live problem for any contact centre using voice biometrics as a factor rather than as a signal. Our notes for contact centres go through that in more detail.
How security teams categorise and handle it
In most operational taxonomies vishing is filed as social engineering, distinguished by delivery channel, and mapped to whatever it enabled: credential theft, initial access, payment diversion, data disclosure. Public adversary-behaviour frameworks catalogue phishing by mechanism and include a voice sub-technique, which is the usual hook for detection and reporting. [VERIFY: framework identifiers and current definitions before citing them]
Simulation is the other place the term appears. Vishing engagements are a standard component of red-team and social-engineering assessments, and they carry ethical and legal weight that email simulations do not: a live call involves a person under pressure, recording a call is regulated differently across jurisdictions, and pretexting as a named third party can create liability. Scoped rules of engagement and an agreed abort phrase are the norm for good reason.
Where audio has been retained, post-incident analysis can sometimes establish whether the speech in a recording was generated. Two cautions belong on any such report. The channel is hostile to the evidence, since telephone coding discards the fine detail an analysis depends on, and phone audio is the hardest condition we work in. And the two possible answers are not equal in weight: a synthetic finding rests on positive evidence, while a human-looking result may only mean the signal was stripped in transit. Reading a result correctly is part of using one responsibly, and our limitations page lists the conditions where we expect to be wrong.
Why there are two pages for one idea
Somebody mid-incident and somebody writing a control framework need different documents. The first needs a procedure they can follow while a phone is in their hand, in words that do not require a security vocabulary. The second needs etymology, taxonomy and classification so that an incident can be recorded and compared.
Splitting them means neither is diluted. The consumer guide carries the callback rule, what the calls sound like and what to do afterwards. This page carries the reference material, and if you arrived here looking for help with a call that has already happened, the first hour after a voice scam is the one to read.
Questions about the term
Is vishing the same thing as voice phishing?
Yes. Vishing is a contraction of the phrase and the two are interchangeable in practice. The contraction is used inside security teams, threat reporting and awareness material; the full phrase is what members of the public search for and understand.
Does vishing require AI voice cloning?
No, and most of it still does not. The technique predates practical cloning by many years and works on pretext, urgency and authority alone. Cloning added a new pretext class rather than replacing the old ones: an attacker can now impersonate a specific known individual rather than a plausible stranger.
Is a cloned executive call vishing or business email compromise?
Both labels are describing different axes, which is the recurring problem with this vocabulary. Vishing names the channel. Business email compromise names the fraud pattern and the target. A cloned call authorising a payment is a voice-channel delivery of the same payment-diversion fraud, and reporting it under only one label loses information.
How is vishing categorised in incident taxonomies?
Generally as social engineering used for initial access or for fraud, sub-classified by delivery channel. Public attack frameworks catalogue phishing by mechanism and include a voice variant. [VERIFY: exact framework technique identifiers and current wording before citing them]
Which controls actually reduce vishing?
Out-of-band verification for any payment or credential change, phishing-resistant authentication that does not depend on a code a person can be talked into reading aloud, a published policy that removes the burden of doubting a senior voice from the individual, and training that teaches a callback procedure rather than a list of audible warning signs.
Practical counterpart
Technical background
Reviewed