
CrowdStrike Shows Attackers Bypass LLM Safety Classifiers by Splitting Harmful Goals Into Benign Subtasks
CrowdStrike says a frontier-model safety classifier blocked about 515 direct bypass attempts, but adversaries can still extract building blocks through genuinely benign subtasks and reassemble working offensive code with an unclassified local model.
Ava Okello, DetectionLondon6 min read
LONDON - Frontier AI models increasingly sit behind a second safety classifier that scores every user request and blocks harmful prompts before the main model answers. CrowdStrike's Cyber Superintelligence Lab reported on Oct. 6 that the strongest publicly deployed classifier it evaluated stopped about 515 distinct direct bypass techniques, including encodings, multi-turn escalation, tokenizer tricks, and novel cognitive-science approaches, with a zero percent direct bypass rate. The same research shows attackers can still circumvent that defense by decomposing an offensive goal into individually benign subtasks, reframing each as legitimate software work, and recomposing the outputs with an unclassified smaller model.
CrowdStrike frames the method as Decompose, Benign Reframe, and Recompose. Figure 1 in the post maps the orchestration pipeline: an unclassified local model breaks a harmful objective into subtasks that are genuinely benign on their own, a classified frontier model answers each reframed request, and the local model stitches the fragments into a working offensive artifact. The firm says an attacker with a free API key and a local open-weight model already has enough to run the pipeline. Validated reframing families include game-modding contexts, detection-engineering prompts, and other dual-use software tasks that naturally use the same code primitives.

In a process-injection case study, three benign requests to the classified frontier model extracted the Windows primitives needed for remote process injection. An unclassified smaller model then assembled a 221-line C program implementing the OpenProcess, VirtualAllocEx, WriteProcessMemory, and CreateRemoteThread chain. A simplified game-trainer framing produced the same result with less orchestration. CrowdStrike also showed that asking the classified model to write Sigma detection rules for a CVE can surface detailed exploitation mechanics because the stated intent is defensive; the smaller model then synthesizes a working proof-of-concept from that knowledge.

Across 10 MITRE ATT&CK-aligned offensive categories, the full pipeline produced working exploit code in nine. The only category that resisted was Defense Evasion aimed at neutralizing runtime security hooks such as EDR inspection. CrowdStrike said the classifier blocked that class at the conceptual level, including prompts that avoided security terminology, because it recognized the idea of patching runtime monitoring functions in memory. The firm notes parallel independent findings in Microsoft Research's Capability Laundering work, which likewise concludes that per-exchange filtering is structurally insufficient when composition happens outside the classifier's observation boundary.

For SOC and detection teams, the operational lesson is architectural rather than a single signature. CrowdStrike argues cross-request semantic accumulation, flagging when an API key's recent requests jointly cover a harmful-composition template, is a natural mitigation but has hard limits: attackers can split queries across providers, rotate keys, or keep orchestration and assembly on local models. The research positions the classifier itself as robust within a per-request threat model and locates the gap in the assumption that per-request evaluation is enough. Defenders who monitor AI coding assistants, developer API usage, and detection-engineering tooling should treat multi-request, multi-model composition as a first-class abuse surface, not only single-prompt jailbreaks.
Sources:
Ava Okello covers detection engineering, EDR telemetry, and SOC hunting for SOCtember from London.
Related stories
Detection
Malware Now Embeds Instructions Meant to Steer AI Analysis Tools, Cisco Talos Finds Across 84 Samples
Threat Intel
Cisco Talos Tracks UAT-11985 Phishing Taiwan Researchers With AI-Assisted Invites and Real-Time Google Login Relays
Threat Intel
Talos Documents CLOSEDQUORUM, Windows Implant That Lets AI Models Vote on C2 Moves
Detection