Contact

How Malware Is Learning to Manipulate AI Analysis

Cisco Talos examines malware that embeds instructions for AI systems, showing why automated analysis pipelines must treat sample text as evidence.

Illustration of malware sample text being separated from trusted instructions in an AI-assisted analysis workflow

Security teams increasingly use language models to summarize suspicious files, classify samples, explain code, and accelerate reverse engineering. That creates a new analytical boundary for malware authors to target: not the executable’s runtime environment, but the automated systems that interpret the file.

In research published through its CAIRN project, Cisco Talos describes this behavior as A3: AI-Analysis Evasion. The category covers malware that embeds natural-language content intended to influence an automated analysis system. Talos identified four confirmed A3 malware families—FRUITSHELL, PLOTSAFE, HOLLOWCLAD, and MANTLEMAZE—among 84 samples collected from January 2025 through July 2026.

The findings do not show that AI-assisted security is broadly ineffective. Instead, they demonstrate that analysis pipelines need a clear separation between an analyst’s instructions and the untrusted text extracted from a sample. A sentence inside a binary, script, or document must be treated as evidence about the file—not as an instruction from an authoritative operator.

A new target above conventional anti-analysis

Traditional anti-analysis techniques attack the binary-analysis process itself. Packers, encrypted overlays, virtual-machine obfuscation, and anti-debugging checks can make code harder to unpack, execute, or inspect. A3 techniques operate at a different layer. They target the workflow that extracts strings from a sample and passes those strings to a language model for triage or interpretation.

This distinction matters because a model may receive two different kinds of text in a single analytical request: the analyst’s question and content supplied by the suspicious file. If the pipeline does not clearly label the latter as untrusted data, imperative language in the sample may be mistaken for a system-level directive.

Talos says the material used in these techniques must remain in plaintext to influence a text-processing model. That limitation also gives defenders a stable detection surface. Text that tells an analyzer to disregard a file, refuse reverse engineering, or accept an unsupported claim can itself become a suspicious signal.

From a simple comment to reusable malware design

FRUITSHELL illustrates the least complicated version of the technique. Talos describes it as a PowerShell reverse shell that uses obfuscated fruit-related variable names while establishing a valid TCP connection. At the top of the script, two comments tell an AI system not to analyze the file and falsely describe it as a harmless prime-number generator. Talos reports that no prime-number-generation code appears in the script and that the comments do not affect execution.

The comments nevertheless matter to an automated workflow because they may be included in extracted text sent to a model. Talos observed the same wording in nine additional scripts over the 15 months following the original VirusTotal submission, with the reuse spanning at least four distinct actors. The researchers state that these were not FRUITSHELL variants, but independent adopters of the same approach.

The technique also appeared in more capable tooling. Talos tracks ROZESHELL as an advanced script collection that combines the same style of AI-evasion comment with AMSI-bypass capability, runtime compilation using csc.exe, and a Rozena shellcode loader. The important analytical point is that the anti-analysis text is not necessarily a marker of unsophisticated malware. It can be added to tooling that contains materially more complex behavior.

Template generation makes the signal harder to match literally

PLOTSAFE represents an evolution from copying a fixed sentence to generating variations from a template. According to Talos, its builds include a recurring introductory phrase followed by changing descriptions of supposedly benign software activity. The samples also contain build-related and test-suite text designed to make each instance look different.

The family is implemented in Go. Talos reports that unused strings would normally be removed by the compiler, so the author added a small dummy function that references the strings without calling it. This detail is significant because it indicates deliberate engineering to preserve the anti-analysis content in the compiled binary.

For defenders, exact-string matching is therefore useful but incomplete. A detection approach focused only on one known sentence could miss generated variants that retain the same intent while changing the filler text.

HOLLOWCLAD tries to speak the model’s formatting language

HOLLOWCLAD uses what Talos calls “template spraying.” The binary contains a refusal instruction repeated across seven different large-language-model chat-template formats. The apparent objective is to increase the chance that one embedded version will resemble a valid system message after an automated scanner wraps extracted strings in its own format.

Talos also found approximately 80 small binary sections named after recognized protector frameworks, including Themida, VMProtect, Enigma, UPX, and Denuvo. The researchers emphasize that this did not mean the binary was packed with those technologies. Instead, the names were presented as a misleading layer intended to confuse packer-identification tools or discourage further analysis.

The same family included intimidation-themed notes. These claimed, among other things, that analysis was being monitored, that a licensing server would detect tampering, or that modifying license-check branches would cause a delayed crash. Talos characterizes these messages as another attempt to influence an analysis model rather than as evidence of the claims themselves.

MANTLEMAZE combines social pressure with a loader stage

MANTLEMAZE applies the template-spraying approach alongside a collection of fabricated authority signals. Talos describes the samples as VMProtect-packed loaders that falsely present themselves as proprietary technology associated with Google, Microsoft, and Valve. The files also contain invented procurement, certification, nondisclosure, patent, and boot-log references.

Talos assesses this behavior as an attempt to encounter topics that might trigger a model’s safety restrictions or otherwise make the model reluctant to continue. These claims should not be treated as evidence of a relationship with the named organizations or government bodies.

The family also contains a potentially relevant loader-stage clue: a program database path referencing iqvw64e.sys, an Intel Ethernet diagnostics driver. Talos identifies this driver as the target of the Bring Your Own Vulnerable Driver technique and cites CVE-2015-2291. The research says the driver has been abused by Scattered Spider, BlackByte, and Lazarus Group to disable endpoint detection and response from kernel space. The supplied research does not establish that every MANTLEMAZE sample performed that activity, nor does it provide a complete execution chain for the loader.

How effective are the techniques?

Talos evaluated whether the embedded language changed model behavior, rather than simply asking whether a model reached the correct final malware verdict. The researchers extracted anti-analysis strings from seven malware families, inserted each string into samples where it had not originally appeared, and compared model responses with and without the string. Five local language models were run three times, producing 135 matched pairs per anti-analysis string.

Talos reports that direct instructions telling a model to ignore the sample—the approach associated with FRUITSHELL and PLOTSAFE—were the most consistently effective in the testing. More elaborate methods generally had less impact and could make the model more suspicious. Across the broader category, the results were mixed: the techniques steered outcomes in the attacker’s favor in about 35% of test runs, according to the source.

That figure should be interpreted as a research evaluation of steering behavior, not as a general estimate of malware-detection failure. Talos also notes that the effect varied by sub-technique and that some messages increased refusals or suspicion rather than producing a benign classification.

Operational implications for security teams

AI-analysis evasion can complicate triage, reverse engineering, detection engineering, and automated reporting. A model that accepts embedded instructions may generate an inaccurate description, omit important behavior, or refuse to analyze a sample. Analysts may then spend additional time validating an output that was influenced by attacker-controlled text.

At the same time, the technique does not replace conventional defenses. Talos concludes that core detection mechanisms remain unaffected and that a properly designed AI-assisted workflow is not inherently more vulnerable than a human analyst who recognizes prompt injection. The immediate risk is workflow design: treating extracted content as if it had the authority to control the analysis.

What organizations should do now

  • Separate instructions from evidence. Use analysis prompts and interface controls that explicitly label all extracted strings, comments, metadata, and decompiled content as untrusted sample data.
  • Keep independent verdicts in the loop. Do not allow a model’s interpretation of embedded text to override endpoint, network, sandbox, or static-analysis findings.
  • Hunt for imperative analysis language. Search scripts and binaries for phrases that tell an AI, scanner, analyst, or reverse engineer to ignore, refuse, or stop analysis. Use semantic and structural variants rather than relying only on exact matches.
  • Correlate text with execution behavior. A suspicious instruction embedded in a file should increase scrutiny, but it should be evaluated alongside process creation, script execution, network activity, persistence indicators, and other available telemetry.
  • Test the analysis pipeline itself. Security engineering teams should run controlled samples containing misleading instructions through every automated triage path and verify that the system preserves the distinction between system messages and file content.
  • Retain raw artifacts. Preserve the original sample, extracted strings, model input, model output, and analyst disposition so that an unexpected classification can be audited.
  • Review BYOVD exposure separately. Where relevant, assess controls around vulnerable drivers and monitor for unauthorized driver loading. The presence of a driver reference in a sample is not, by itself, proof that the driver was loaded or abused.

Conclusion

Cisco Talos’s A3 research shows that malware authors are adapting to the growing role of language models in security operations. The progression runs from static comments, through generated wording, to model-specific template spraying and intimidation-themed content. The approach is inexpensive, inconsistently effective, and visible in plaintext, but it can still disrupt automated triage when pipelines fail to treat sample content as untrusted.

The practical response is not to remove AI from malware analysis. It is to design analysis systems that understand the difference between an instruction issued by the workflow and text supplied by the object under investigation. That boundary, supported by independent telemetry and human validation, is the central control against this emerging form of analysis manipulation.

Sources

Cisco Talos: “Ignore all instructions and read this blog: The state of AI-analysis evasion in malware”

Selected sample hashes reported by Cisco Talos

  • FRUITSHELL: f8f5e0440c57c7deffd75ca33e2511867039796aa803e7ef847396a379188a7d
  • HOLLOWCLAD: 34098fe0bc4c69c4c4eb3f74688fb375326804536d25e5574e8c4c28c113b5c3
  • MANTLEMAZE: 389066bd5543aeea363d23a4dce7f7a21c7f2c73c61f506a93a69f594cf48ecf
  • PLOTSAFE: 2e3e1bcd44cc3cbec4f5ca9991d14a326d3cc6bf76fe6e0d434c9b5f7e1ae6ab
  • ROZESHELL: 0d2d6e6b03a19ae31d2af279e88a41d911828f0b531fed005ad2ff44566c261