Halo: A Real-Time Anti-Deepfake Tool That Detects AI Face-Swapping in Video Calls

Halo detects AI face-swapping in real-time during video calls to prevent deepfake fraud.
Halo is an on-device deepfake detection tool that analyzes video feeds in real time during Zoom, Teams, and Google Meet calls to identify AI-synthesized faces. Built by Scam AI, it processes all detection locally to protect privacy, targeting high-risk scenarios like remote hiring fraud and fake executive wire transfer scams. The tool faces challenges in balancing accuracy, computational overhead, and robustness against evolving forgery techniques.
Video Calls Are Becoming the New Deepfake Battleground
The person on the other side of your next video meeting might not be real. This isn't science fiction—it's a scam technique that's actively happening worldwide. Recently, a deepfake detection product called Halo launched on Product Hunt, earning 108 upvotes and ranking 6th on the daily leaderboard with its core value proposition of "real-time detection of synthetic faces in video calls." It was categorized under Meetings, Artificial Intelligence, and Security.
As deepfake technology becomes increasingly accessible, attackers can now impersonate others using real-time synthetic faces on mainstream meeting platforms like Zoom, Teams, and Google Meet. Deepfake technology originated around 2017 with GAN-based (Generative Adversarial Network) face-swapping algorithms. The core principle involves using encoder-decoder architectures or diffusion models to map one person's facial features onto another's face in real time. A GAN consists of two neural networks—a generator that synthesizes the most realistic face images possible, and a discriminator that distinguishes real from fake. Through adversarial training, both continuously improve until the generated synthetic faces become indistinguishable to the human eye. In recent years, the rise of Diffusion Models has further improved image generation quality by generating images through a gradual denoising process, surpassing traditional GAN methods in detail fidelity and diversity. Early face-swapping required thousands of training images and days of model training time. Today, with pre-trained large models and open-source tools (such as DeepFaceLive and FaceFusion), attackers need only a few photos or a short video of their target to deploy in minutes, injecting the synthetic feed into video conferencing software through virtual cameras (like OBS Virtual Camera)—which meeting platforms treat as ordinary camera signals with no way to distinguish them. This democratization of technology has caused identity fraud risk in video calls to increase exponentially.
From faking CEO wire transfer instructions to impersonating job candidates in remote interviews, these face-swapping scams routinely cause losses of hundreds of thousands or even millions of dollars. The vast majority of ordinary users—and even corporate employees—have absolutely no means to distinguish real from fake on the other side of their screen.

What Is Halo: An On-Device Real-Time Deepfake Detection Shield
Built by the Scam AI team, Halo has a very focused product positioning: detecting and flagging AI-synthesized faces in real time during video calls. It works by analyzing video frames in the background while you're in a Zoom, Teams, or Google Meet meeting, immediately alerting you whenever it detects a face that appears to be AI-generated or face-swapped.
The product makes one critical technical promise—all detection is performed entirely on your device. This means your meeting footage never needs to be uploaded to cloud servers for analysis, reducing both the risk of privacy breaches and compliance concerns about handing sensitive business meeting content to third parties.
From a technical implementation perspective, on-device AI inference has made significant strides in recent years. Thanks to model quantization (compressing 32-bit floating-point parameters to 8-bit or even 4-bit integers, shrinking model size by 4-8x and improving inference speed by 2-4x with minimal accuracy loss), knowledge distillation (having a much smaller "student model" learn the output behavior of a large "teacher model," retaining most detection capability while drastically reducing computation), and Neural Architecture Search (NAS, using automated search algorithms to find optimal accuracy-efficiency combinations in vast architecture spaces), deep learning models that previously required cloud GPU clusters can now achieve near-real-time inference on consumer laptops equipped with Apple Silicon (M-series chips with 16-core Neural Engines delivering up to 15.8 TOPS), Intel NPUs (like the neural processing unit integrated in Meteor Lake), or discrete GPUs. Inference frameworks like Apple's Core ML, Intel's OpenVINO, and NVIDIA's TensorRT further optimize local deployment efficiency through graph optimization, operator fusion, and hardware-specific instruction set acceleration. This makes it possible for products like Halo to achieve fully offline privacy protection without sacrificing too much detection accuracy.
For privacy-sensitive scenarios in finance, law, and recruiting, this local-first design is crucial. Globally, the EU's General Data Protection Regulation (GDPR) explicitly classifies facial data as "special category personal data," requiring stricter legal bases for processing. Illinois' Biometric Information Privacy Act (BIPA) requires companies to obtain written informed consent before collecting facial geometry data, with statutory damages of $1,000-$5,000 per violation (multiple class-action lawsuits have resulted in settlements of hundreds of millions of dollars). China's Personal Information Protection Law classifies facial recognition information as "sensitive personal information," requiring separate consent and impact assessments before processing. If a detection tool needed to upload video frames to the cloud for analysis, enterprises would face additional burdens including cross-border data transfer compliance reviews (such as Standard Contractual Clauses or adequacy determinations under GDPR), third-party data processing agreements, and potential violations of labor law provisions if employees haven't explicitly consented. Halo's purely local processing architecture eliminates these legal risks at the source—data never leaves the user's device, dramatically reducing compliance complexity.
Two High-Risk Scenarios Halo Directly Addresses
Halo's team specifically highlights two typical use cases:
-
Identity verification in remote hiring: Cases of job candidates using deepfake face-swapping to impersonate others during interviews are increasingly common, especially in remote outsourcing for technical roles. Halo helps HR confirm candidates' true identities. According to a 2022 Public Service Announcement (PSA) from the FBI's Internet Crime Complaint Center (IC3), complaints about using deepfake technology to impersonate others in remote interviews have surged, particularly concentrated in high-paying remote positions in IT, software development, and database management. Attackers are typically motivated by gaining access to internal enterprise systems or collecting salaries fraudulently—once successfully hired, these imposters may access core assets including customer personally identifiable information (PII), financial records, corporate code repositories, and intellectual property. More seriously, some cases are linked to nation-state threat actors; North Korean IT workers using stolen or fabricated identities to gain remote employment at Western tech companies have been subject to multiple sanctions and warnings from the U.S. Treasury Department. This elevates identity verification in remote hiring from an HR process issue to a national security concern.
-
Confirming wire transfer counterparties: Finance staff receiving video instructions from "executives" demanding urgent transfers represents an upgraded version of Business Email Compromise (BEC) scams, often with even more devastating losses. Traditional BEC relies on spoofed email addresses or compromised executive accounts to issue fraudulent transfer instructions, while video-based BEC uses real-time deepfake face-swapping to create the trust of "seeing it with your own eyes," completely demolishing the long-held security assumption that "if you see the real person, it's trustworthy." In early 2024, a finance employee at a multinational company in Hong Kong (reportedly the Hong Kong branch of British engineering group Arup) was instructed to transfer funds by a deepfaked "CFO" and multiple "colleagues" during a multi-participant video conference, ultimately losing HK$200 million (approximately US$25.6 million). According to Hong Kong police, every participant in the meeting except the victim was an AI-synthesized fake persona. This case marked the escalation of deepfake scams from single-target attacks to multi-character coordinated deception, demonstrating that human eyesight alone is completely inadequate against this threat level. Halo can help verify counterparty identity before funds are transferred.
Halo's tagline "Catches it before it costs you" targets precisely these pain points—it aims to sound the alarm before money flows out or a bad hire is made.
Why Real-Time Detection Matters More Than Post-Hoc Forensics
Deepfake detection isn't a new concept, but embedding deepfake detection capability into real-time video call workflows remains a relatively emerging and high-demand direction. Previous anti-deepfake tools were primarily used for post-hoc forensics—analyzing whether a pre-recorded video had been tampered with. These tools typically rely on multiple technical approaches: frequency domain analysis performs Fourier or wavelet transforms on images to detect spectral artifacts unique to GAN-generated images (GAN upsampling operations leave periodic checkerboard-pattern artifacts in the high-frequency domain); physiological signal detection uses remote photoplethysmography (rPPG) technology to extract heartbeat signals by analyzing minute periodic color changes in facial skin—real faces exhibit pulse signals consistent with heart rate, while synthetic faces typically lack this physiological characteristic; temporal consistency analysis checks whether lighting direction, shadow casting angles, and facial specular highlights remain physically consistent between adjacent video frames. These methods can achieve high detection accuracy in offline environments (academic papers commonly report AUC values above 95%), but processing a video segment often takes seconds to minutes.
Halo's core value lies in "real-time capability": scams often happen in a flash, and post-hoc detection cannot recover losses already incurred. Only by issuing warnings during the call itself can a scam truly be intercepted. The technical challenges of real-time detection far exceed those of offline analysis—it requires completing the entire pipeline of face detection, feature extraction, and authenticity classification within a 24-30 fps video stream, while keeping end-to-end latency under 100 milliseconds to avoid impacting user experience. This means each frame has a processing budget of only about 33 milliseconds, and the complete detection pipeline typically includes face localization (using lightweight models like BlazeFace or SCRFD), face alignment and cropping, feature extraction (extracting discriminative features via convolutional neural networks or Vision Transformers), and classification decisions (outputting real/fake probability values)—each step needing to complete at the millisecond level. This requires a delicate balance between accuracy and speed in the detection model.
From an industry trend perspective, as generative AI becomes more widespread and real-time face-swapping tools (such as various face-swap live streaming plugins) grow more capable, detection and forgery are effectively locked in a continuous "arms race." The essence of this arms race is an ever-escalating adversarial process: forgers continuously optimize generation models to eliminate detectable artifacts (such as improving Poisson blending algorithms at facial boundaries, using super-resolution networks to fix texture inconsistencies, and matching environmental reflections through lighting estimation models), while detectors must continuously discover deeper, harder-to-forge discriminative features. Current cutting-edge academic directions include: leveraging the physical consistency of facial muscle micro-action units (Action Units) defined in the Facial Action Coding System (FACS)—real faces follow anatomical constraints in AU combinations, while synthetic faces often produce unnatural AU co-occurrence patterns; verifying 3D geometric plausibility through ray tracing—analyzing whether specular highlight points on pupils, nose tips, cheeks, and other areas point toward the same light source; and capturing global spatiotemporal artifacts using attention-based Vision Transformer architectures—Transformer self-attention mechanisms can model long-range dependencies within and between video frames, capturing global inconsistencies imperceptible to the human eye. This also means that the biggest challenge for real-time detection products like Halo is the continuous evolution capability of detection models—as forgery technology iterates, detection algorithms must keep pace or stay ahead.
Technical Challenges and Unverified Questions Facing Halo
As a newly launched product, Halo still has several key questions awaiting market validation:
-
Detection accuracy and false positive rate: Excessive false positives disrupt normal meetings, while low recall rates render the tool useless. Real people under low-light, low-quality, or network-lag conditions may be misidentified as AI-synthesized. In machine learning terms, this involves the tradeoff between Precision (the proportion of samples flagged as "fake" that are actually fake) and Recall (the proportion of actual fake samples that are successfully detected). For security products, the ideal state maintains extremely low false positive rates (high precision) while sustaining high detection rates (high recall), but in reality these often work against each other, requiring reasonable decision thresholds based on the application scenario. Video compression artifacts (such as blocking effects and mosquito noise from DCT transforms in H.264/H.265 encoding), screen freezes and pixelation from network packet loss, and random noise and color shifts from low-end camera sensors can all introduce image degradation visually similar to deepfake artifacts, significantly increasing misclassification risk. Maintaining robust detection performance under degraded conditions is one of the most challenging engineering problems in practical deployment.
-
Local computational overhead: Real-time frame-by-frame analysis places certain demands on device performance, and whether it runs smoothly on ordinary office laptops requires real-world testing. At 30fps video stream, the detection model needs to complete full single-frame inference within approximately 33 milliseconds while not over-consuming CPU/GPU resources to the point of affecting the video conferencing software itself—after all, users still need to simultaneously run Zoom or Teams clients, screen sharing, and other office applications. This places high demands on lightweight model design (using mobile-optimized architectures like MobileNet or EfficientNet) and inference pipeline engineering optimization (tradeoffs like batch processing strategies, asynchronous inference, and keyframe sampling rather than processing every frame). One possible compromise is a tiered detection strategy: first using an extremely lightweight model for rapid suspicious frame screening, then triggering more detailed deep detection only for suspicious frames, thereby balancing overall computational consumption and detection accuracy.
-
Robustness against novel forgery techniques: Whether Halo can maintain its detection advantage against continuously evolving deepfake methods is a long-term proposition. Particularly once attackers become aware of the detection tool's existence, they may specifically launch adversarial attacks against the detection algorithm. This is a core threat in machine learning security—attackers can cause detection models to classify forged faces as real by adding imperceptible perturbations to pixel values in the generated image (typically with L∞ norm within 4/255). Academic research has demonstrated that most deep learning-based classifiers can see detection accuracy plummet from above 95% to below 10% when facing carefully crafted adversarial examples. This requires the product team to establish continuous model updating and adversarial training mechanisms—intentionally injecting adversarial examples during training to enhance model robustness while maintaining rapid response capability against new forgery methods.
Conclusion: An Essential Security Tool in the Age of AI Trust Crisis
Halo's emergence is essentially a microcosm of AI's double-edged sword effect—generative AI brings both an efficiency revolution and an unprecedented trust crisis. When "seeing is no longer believing," a deepfake detection tool that can answer "is the person on the other side real?" during video calls is transitioning from a nice-to-have to an essential requirement in high-risk scenarios.
The impact of this trust crisis extends beyond individual fraud cases. At the enterprise level, Gartner predicts that by 2026, 30% of enterprises will consider existing identity verification and authentication solutions unreliable when used alone, precisely because deepfake technology systematically undermines traditional identity verification methods (including video interviews, voice confirmation, etc.). This prediction is driving more enterprises to evaluate multi-modal identity verification solutions, incorporating deepfake detection as a new link in the identity verification chain. Meanwhile, platform providers like Microsoft and Zoom are actively exploring content authenticity verification standards—for example, the C2PA (Coalition for Content Provenance and Authenticity) alliance, co-founded by Adobe, Microsoft, BBC, and others, is promoting content provenance authentication protocols that embed cryptographically signed metadata in media files to prove content creation sources and editing history. However, the C2PA standard requires comprehensive support across the entire content production and distribution chain (from camera hardware, operating systems, and editing software to publishing platforms), and its full deployment in real-time video call scenarios will still take years. During this transition period, third-party real-time detection tools like Halo fill a critical security gap, providing users with a trust verification layer independent of the platform.
For enterprise security teams, finance staff, HR professionals, and anyone who needs to establish trust with strangers via video, the "real-time anti-forgery" direction that Halo represents deserves ongoing attention. Whether it can deliver on its promise of "catching scams before losses occur" ultimately depends on its detection performance in real-world environments. But at the very least, it raises a question that increasingly cannot be ignored: Are you sure the person on the other side of your screen is real?
Related articles

Generate Xianxia Wallpapers with a Single Sentence: A Hands-On Comparison of Three AI Agents with Skill Enhancement
Testing the "Eastern Xianxia Visual Director" Skill across Codex, WorkBody, and Grog to see how a single plain sentence becomes stunning xianxia wallpaper art.

Pony Language: The Lock-Free Concurrent and Memory-Safe Programming Language Is Still Alive
Pony is an Actor model-based programming language that guarantees memory safety and data-race freedom at compile time through its reference capabilities system.

Google AI Marketing Tools Explained: A Deep Dive into Google Ads and Analytics Agentic Features
Google introduces new AI agentic features in Google Ads and Analytics for automated ad optimization and proactive data insights. A deep dive into how these tools reshape digital marketing workflows.