WhatsApp is testing Scam Alert, an optional feature that identifies possible scam messages while protecting user privacy. The tool uses an on-device machine-learning model to review messages without sending chat content elsewhere.
However, the feature does not automatically share messages with WhatsApp, Meta, or other third parties. Instead, the model runs directly on a user’s device and checks conversations for scam-related patterns.
The feature is currently in a limited beta rollout. Additionally, WhatsApp is sharing technical details early so security researchers can examine the system and find possible weaknesses before a wider release.
On-device scam detection and user control
Once users enable Scam Alert, their device downloads a machine-learning model. The model checks incoming messages from people who are not saved contacts. It looks for language patterns and conversation structures linked to known scams.
Moreover, the model learns from patterns found in scam conversations that users previously reported. Therefore, it can identify signals without needing access to complete chat histories.
When the system detects a possible scam, users receive a warning inside the chat. The warning appears only to the user. They can then block or report the sender, or continue the conversation.
If the warning is incorrect, users can mark the chat as trusted. After that, Scam Alert will stop flagging that conversation. Users can also choose to share their last five messages after marking a chat as trusted, which helps improve the feature.
Privacy-focused measurements and security checks
WhatsApp said the system follows three main principles: on-device processing, no automatic reporting, and user control.
Furthermore, all message analysis happens on the device. WhatsApp cannot start sharing user information unless the user chooses to report a message.
The system still needs measurements to understand how well Scam Alert works. Therefore, it uses aggregated information instead of collecting message content.
Warning counts show how often the model creates scam alerts. User-action counts show what people do after receiving those alerts. For example, users may trust a chat or block and report a sender.
Meanwhile, confidential federated analytics protects these limited measurements. The device converts information into aggregate counts before sending any data.
Raw signals remain on the device and are removed after a set retention period. Additionally, encrypted transfers and secure computing environments protect the collected statistics.
The system combines data from multiple devices before adding further privacy protections. Differential privacy adds controlled changes to the data, while minimum group requirements prevent results from small groups.
Moreover, an Oblivious HTTP relay removes IP addresses from requests. Anonymous credentials also confirm that requests come from genuine WhatsApp clients without identifying individual devices.
WhatsApp also prevents itself from sending a specific model to a specific user. Instead, models come from a content delivery network, which allows updates as scam methods change.
Transparency for users and researchers
WhatsApp is adding transparency tools to Scam Alert. Users can enable logs that show analysed messages, warnings, model versions, and related activity details.
Furthermore, users can access these records through Scam Alert Activity under Request Info. The feature also supports security research through an expanded bug bounty programme.
Researchers can examine whether message content stays on the device and whether the model works as intended. Additionally, they can access model weights to test how the system responds to different inputs.
Finally, these checks aim to improve trust while helping users understand how Scam Alert identifies potential fraud.








