How Deepfake Attacks Are Evolving Faster Than Defenses: What Enterprises Need to Know
Deepfake attacks are no longer just about creating convincing fake videos; they're now sophisticated social engineering weapons designed to manipulate employees into authorizing payments, changing payroll, or resetting credentials. Unlike earlier deepfakes that relied on visual trickery alone, modern attacks combine cloned voices, synthetic video calls, and coordinated messaging across multiple channels to impersonate trusted executives and colleagues.
What Makes Today's Deepfake Attacks Different From Earlier Synthetic Media?
The distinction between a deepfake and other altered media matters because detection methods differ significantly. A deepfake is synthetic content created with artificial intelligence to make a person appear to say or do something that never happened, specifically designed to misrepresent identity or intent. This differs from a "shallowfake," which uses simpler techniques like cutting video out of context or splicing unrelated statements without advanced AI. A digital injection attack, meanwhile, inserts fabricated audio or video directly into a live communication or authentication process, making it appear authentic to both people and systems that assume the camera feed is genuine.
The defining risk is false authenticity. When a target receives a message that appears to come from a chief financial officer, the message becomes far more persuasive if an accompanying voice note sounds exactly like the CFO. A video call that appears to show the CFO giving instructions strengthens the deception even further. This coordinated pattern is called multimodal impersonation, in which a cyberattacker uses two or more communication formats, identities, or channels to make the same false request appear credible.
How Are Attackers Using Audio-Visual Deepfakes to Target Business Workflows?
Cyberattackers are deliberately targeting the workflow rather than the technology itself. The highest-risk areas include payments, payroll changes, help-desk password resets, employee onboarding, recruiting, and executive communications. A 2024 incident involving a Hong Kong employee demonstrates the real-world danger. An employee received a deepfake video call that appeared to show a company executive authorizing a fraudulent transfer. The result was an approximately $25 million loss, showing how convincing media becomes dangerous once it overrides process controls.
Vishing, or voice phishing, conducted by phone or voicemail, has become particularly effective with AI voice cloning. A cyberattacker can create a convincing imitation of a manager or family member, often followed by a request to disclose a one-time code or authorize a payment. Smishing, which is phishing delivered by SMS or messaging services, can direct a recipient to a fake login page and then use a cloned voice call to pressure the recipient into completing the action.
Attackers also use open-source intelligence, or OSINT, to personalize their deepfake attacks. They gather information from company websites, professional profiles, conference recordings, social media, job postings, and public filings to identify reporting lines, speech patterns, travel schedules, vendors, and approval workflows. This allows them to craft highly targeted requests instead of sending obviously generic messages.
How to Protect Your Organization Against Deepfake Attacks
- Independent Callback Verification: Security teams should teach employees to verify the request through a trusted channel rather than judging the media alone. A familiar face or voice is one signal, but identity requires independent proof. When a high-value request arrives, pause and call the person back using a known phone number or email address to confirm the request is legitimate.
- Dual Approval and Workflow Governance: Implement separation of duties so that no single person can authorize high-value transactions. Require multiple approvals for payments, payroll changes, and credential resets, ensuring that even if one person is deceived by a deepfake, a second reviewer can catch the fraud.
- Phishing-Resistant Multi-Factor Authentication: Deploy MFA methods that cannot be bypassed by social engineering, such as hardware security keys or biometric verification. Standard SMS-based codes can be intercepted or manipulated, but phishing-resistant MFA adds a technical barrier that deepfakes alone cannot overcome.
- Biometrics and Liveness Checks: Biometrics and liveness checks can confirm continuity with an enrolled person, yet they should never authorize a high-value action by themselves. Use them as one layer in a multi-layered defense, not as a standalone control.
- Employee Behavior Change Through Phishing Simulations: Measured behavior change through multi-channel phishing simulations turns employee judgment into the most reliable deepfake control an organization holds. Regular training and simulated attacks help employees recognize social engineering tactics and respond appropriately.
Media inspection alone is unreliable. A convincing impersonation can fool visual review, which is why the FBI and FTC both recommend the same response: pause high-risk requests and verify them through an independent trusted channel.
Why Multimodal AI Is Expanding Beyond Entertainment Into Enterprise Security
While deepfakes represent a misuse of audio-visual AI, legitimate multimodal AI is simultaneously transforming how enterprises operate. Multimodal AI refers to systems that can process and understand multiple types of input, such as text, images, audio, and video, in a single conversation or workflow. This technology is now being integrated into smart eyewear and enterprise applications to improve productivity and accessibility.
Innovative Eyewear announced a major update to its Lucyd app that includes multimodal AI conversations, allowing users to move seamlessly between visual and audio conversations while preserving context. Users can start a conversation in visual mode on their phone and continue it in audio mode through their smart glasses, with the AI maintaining full conversational context. This flexibility demonstrates how audio-visual AI can enhance legitimate workflows, but it also underscores why organizations must distinguish between beneficial multimodal tools and malicious deepfake attacks.
The audiovisual archiving sector is also exploring multimodal AI for legitimate purposes. At the FIAT/IFTA World Conference 2025 in Rome, researchers presented work on multimodal retrieval approaches for film archives. The RTVE Archive, which holds extensive audiovisual collections from Spanish television dating back to the 1950s, is using multimodal AI to search news film collections that lack accompanying sound or informative titles. By embedding both textual queries and video visual features into a shared vector space using the CLIP model, researchers can query archives using natural language or example images, making historical content more accessible.
The contrast is stark: the same audio-visual AI capabilities that enable deepfake attacks can also restore historical archives, improve accessibility for people with disabilities, and streamline business communications. The difference lies in intent, transparency, and the controls surrounding deployment.
As deepfake technology becomes more sophisticated, organizations cannot rely on employees to spot fakes through visual inspection alone. The most effective defense combines technical controls, workflow governance, and employee training to verify high-risk requests through independent channels. The stakes are clear: a single convincing deepfake can cost millions.