Executive Summary
On September 22, 2026, Microsoft published an analysis of EvilTokens alongside news that its Digital Crimes Unit had facilitated a coordinated disruption of the service's infrastructure. EvilTokens was a phishing-as-a-service platform that appeared in February 2026, sold for $1,500 up front and $500 per month, and specialised in abusing the OAuth device authorization flow to steal session tokens without ever touching a password. Microsoft tracks the operator as Storm-2992 and associates the platform with more than 12,000 compromised inboxes across more than 10,000 organisations.
PhishPond has covered this family before — EvilTokens surfaced in our Kali365 analysis as one of the AI-generated device-code lookalikes advertised on Telegram. The takedown is a genuine win and a temporary one. The device authorization grant still exists, still behaves exactly as RFC 8628 specifies, and still issues tokens to whichever client is polling. What is worth studying is not the service's removal but its architecture, because the design decisions that made EvilTokens effective are cheap to reproduce and are already visible in adjacent kits.
AttackAttack Overview
Device code phishing abuses a legitimate OAuth flow built for devices with constrained input — smart TVs, CLI tools, IoT consoles. The device asks the identity provider for a code, displays it, and polls until a human approves it elsewhere. Nothing in that exchange validates that the human approving the code and the device requesting it are the same party. That gap is the whole attack.
What EvilTokens sold was not the technique. The technique is free and documented in an IETF RFC. What it sold was the operational scaffolding around it: lure generation, hosting that does not get blocked, a polling loop that does not drop tokens, and a dashboard that keeps stolen sessions alive after the victim closes their laptop. Understanding the product helps calibrate what the takedown removed.
AttackOperator Workflow
The customer-facing flow is tight. A victim receives a lure with a malicious URL or attachment. A background script on the landing page generates a live device code against the Microsoft identity provider in real time — not a static code baked into the page, which is what keeps the code inside its validity window. The victim sees the code presented with "Copy Code" and "Continue with Microsoft" buttons, and an automated script copies the code to the clipboard to reduce friction and transcription errors.
The script then enters a polling state, pinging the threat actor's `/state` endpoint every three to five seconds. While the victim has not yet completed authentication, polling returns "pending." When the victim enters the code on the genuine `microsoft.com/devicelogin` page and approves, the operator's client receives tokens. The victim's credentials were never exposed, never transited the phishing page, and never appeared in any log as stolen — because they were not.
Within roughly ten minutes of a successful issue, operators often register a new device against the tenant to obtain a Primary Refresh Token. That step converts a session that will expire into access that persists. It is the single most consequential action in the chain and the one most worth alerting on.
AttackInfrastructure & Tradecraft
EvilTokens shipped as a product, and the component list reads like one: a Telegram bot handling distribution and customer support, a control panel with deployment options, 44 pre-built email templates with an AI-powered writing assistant, a token management dashboard with auto-refresh, Microsoft Graph reconnaissance tooling for mapping the victim organisation, and inbox keyword scanning that pushed alerts back to the operator over Telegram.
The hosting choices are the part defenders should study hardest. Customers could deploy on Cloudflare Workers, Bunny, or conventional PHP hosting, and the platform leaned heavily on serverless infrastructure — Vercel (`*.vercel.app`), Cloudflare Workers (`*.workers.dev`) and AWS Lambda. Multi-stage redirection ran through compromised legitimate domains so that campaign traffic blended with ordinary enterprise cloud activity.
This is the design decision most likely to be inherited by whatever replaces EvilTokens, and it is worth being precise about why it works. A `*.workers.dev` or `*.vercel.app` hostname arrives with a valid certificate, a reputable parent domain, infrastructure your users legitimately reach every day, and effectively zero cost to abandon and replace. Domain-reputation blocking has very little to grip. Blocking the parent domains wholesale is not viable for most organisations because real engineering teams depend on them. The same ANY.RUN-observed pattern shows up across unrelated campaigns this quarter, which suggests convergent evolution rather than shared code.
Evasion on the delivery side was conventional but well executed: fake CAPTCHA verification pages, image-based links that defeat text extraction, and multi-stage delivery through attachments.
Post-compromise, the platform's own tooling took over. Operators reviewed victim mailboxes with AI filtering to identify high-value targets, created malicious inbox rules for persistence and to conceal their own correspondence, ran Graph reconnaissance to map the organisation, registered devices for durable access, and pursued targeted financial exfiltration from users with wire transfer authority. Victim concentration was highest in the United States, Canada, the United Kingdom, Australia, India and France, across wholesale distribution, construction, financial services, real estate, higher education and healthcare.
DetectionDetection Opportunities
The defining property of this attack — no credential theft — determines where the telemetry is. There is no stolen-password event because no password was stolen. There is no lookalike-domain hit on the authentication step because authentication happened on Microsoft's real domain. Everything useful lives in the sign-in and audit logs after the fact.
Microsoft maps detections across the chain: anomalous OAuth device code authentication activity at initial access, anomalous token exchange following device code authentication at credential access, suspicious Entra device join or registration at persistence, anomalous Microsoft Graph API request volumes at discovery, and suspicious inbox rule creation post-authentication at defense evasion.
The base signal remains the authentication protocol itself. In most tenants, interactive device-code sign-ins should be rare and attributable to a known set of developer tooling.
SigninLogs
| where TimeGenerated > ago(24h)
| where AuthenticationProtocol == "deviceCode"
| where ResultType == 0
| summarize SigninCount = count(), Apps = make_set(AppDisplayName),
Resources = make_set(ResourceDisplayName), IPs = make_set(IPAddress)
by UserPrincipalName, bin(TimeGenerated, 1h)
| where Apps !in (dynamic(["Azure CLI", "Visual Studio Code"]))The higher-value rule is the one that targets the ten-minute PRT grab, because that step is what converts a transient theft into durable access. Join device-code sign-ins against directory audit events for device registration inside a short window.
let window = 15m;
let deviceCodeSignins =
SigninLogs
| where TimeGenerated > ago(24h)
| where AuthenticationProtocol == "deviceCode" and ResultType == 0
| project SigninTime = TimeGenerated, UserPrincipalName,
SigninIP = IPAddress, AppDisplayName;
AuditLogs
| where TimeGenerated > ago(24h)
| where OperationName in ("Add device", "Add registered owner to device",
"Register device", "Add delegated permission grant",
"Consent to application")
| extend Actor = tostring(InitiatedBy.user.userPrincipalName)
| join kind=inner deviceCodeSignins on $left.Actor == $right.UserPrincipalName
| where TimeGenerated between (SigninTime .. SigninTime + window)
| project SigninTime, AuditTime = TimeGenerated, Actor, OperationName,
AppDisplayName, SigninIP, ResultInbox rule creation shortly after a device-code sign-in is the third leg, and it is the one that most reliably indicates the operator has moved from access to objective. Alert on new rules that forward externally, delete on receipt, or move messages matching finance keywords into low-traffic folders, correlated against a device-code sign-in in the preceding hours.
ValidationValidation Workflow
Prove each leg separately in an authorized test tenant with disposable accounts and no production data.
Generate the base signal honestly: run `az login --use-device-code` as a test user and complete the prompt on a second device. Confirm a `SigninLogs` event appears with `AuthenticationProtocol` equal to `deviceCode`, and record the `AppDisplayName` value your legitimate tooling produces. That value is what belongs on your allow-list, and discovering it empirically beats guessing it.
Then validate the correlation rule, which is the one that matters. After the device-code sign-in completes, register a device against the test tenant as that same test user and confirm the join fires inside the fifteen-minute window. Vary the delay — register at twenty minutes instead of five — and confirm the rule goes quiet, so you know exactly what your window is and is not covering. Tune it against your own observed token and registration timing rather than inheriting fifteen minutes from this article.
Finally, confirm your inbox-rule alerting actually fires on the rule shapes that matter. Create a test rule that forwards externally and one that moves finance-keyword mail to a rarely-opened folder, and verify both surface. Rules that only alert on external forwarding miss the concealment pattern EvilTokens customers were trained to use.
GapsEvasion & Gaps
The device-code allow-list is the obvious weak point, and it is unavoidable. Organisations with real developer use of the flow must maintain exceptions, and an operator who selects a first-party client identifier that your engineers already use lands inside the exception. Keep the list minimal and add a second condition: alert when an allow-listed client reaches a resource it has no business reaching.
Serverless hosting defeats most infrastructure-side detection by design. You cannot usefully block `*.workers.dev` or `*.vercel.app`, the certificates are valid, and each campaign's hostname is disposable. Any detection strategy anchored to the delivery infrastructure will decay faster than the kit does. This is the structural reason the detections above all sit in identity telemetry rather than network or mail telemetry.
The fifteen-minute correlation window is a tuning parameter and an evasion seam simultaneously. Microsoft's observation is that registration "often" happens within ten minutes, which is a statement about convenience rather than constraint. An operator with patience can wait out any window you set. The base device-code alert is what covers that case, which is the argument for keeping it even when it is noisy.
Finally, the takedown itself creates a measurement gap. Storm-2992's infrastructure being disrupted does not retire the 12,000 already-compromised inboxes, the inbox rules already planted in them, or the devices already registered. Organisations that were victims before September 22 are not remediated by the disruption, and the natural human response to a takedown headline — treating the threat as closed — is the gap most likely to be exploited this quarter.
Defensive Recommendations
Block the device code flow wherever possible. Microsoft's own first recommendation is also the only control that removes the attack surface rather than observing it. Scope a Conditional Access policy that blocks the device code grant for every user who does not have a documented engineering need, and review the exception group quarterly.
Require phishing-resistant MFA — FIDO2 tokens, or Microsoft Authenticator with passkey — and pair it with the understanding, covered in our analysis of passkey-themed social engineering, that the enrollment and recovery paths around those methods are now themselves under attack.
Implement sign-in risk policies with Conditional Access, configure anti-phishing policies and Safe Links in Defender for Office 365, and enable zero-hour auto purge so that lures delivered before a signature existed still get pulled from mailboxes.
Build the revocation playbook before you need it, and get the order right. Revoke user refresh tokens with the `revokeSignInSessions` API call, then disable the compromised account for containment, then audit the preceding twenty-four hours of audit logs for device registrations, app consents and inbox rules, and reverse each one. Disabling the account first leaves issued tokens valid, which is the failure mode that has produced most of the public incidents tied to this flow.
Then hunt retrospectively. If your organisation was in scope between February and September 2026 — particularly in wholesale distribution, construction, financial services, real estate, higher education or healthcare — assume the disruption did not clean anything up on your side. Look for devices registered within minutes of a device-code sign-in, inbox rules created by accounts that have no history of creating them, and Graph reconnaissance volume from identities that should not be generating any. The service is gone. The access it sold is not.