Attacks on AI

Attacks against AI systems: models, agents, the applications built on them and their supply chain. Each entry is a short summary with a direct link to the original report or article.

52 entries, last updated 21 Sep 2026, 10:18 UTC. All entries.

September 2026

  1. Gen DigitalVendor report

    Infostealers Have Found a New Target: Your AI Agent (opens gendigital.com)

    Gen Digital's telemetry shows infostealers like Amatera, Remus, CallbackBeaver, and Djinn Stealer have added AI coding agents (Claude, Cursor, Codex, Cline, OpenCode) to their collection rules, harvesting tokens, MCP credentials, and prompt histories. This expands the infostealer economy to target local AI agent data as a new high-value asset alongside browser and wallet credentials.

    Category
    AI-Targeted
    Malware
    Amatera, Remus, CallbackBeaver, BeeStealer, STG Stealer, HydraStealer, APEX Stealer, Otter Stealer

August 2026

  1. HuntressVendor report

    The AI Attack Surface: How Threat Actors Abuse Trusted AI Platforms (opens huntress.com)

    Huntress documents campaigns abusing legitimate AI platform features, Claude Artifacts, claude.ai/share links, and shared ChatGPT/Grok conversations, to host phishing and ClickFix-style lures on trusted domains, leading victims to install SectopRAT, MacSync stealer, or AMOS stealer. These attacks exploit trust in AI branding and domains combined with SEO/malvertising rather than flaws in the AI models themselves, hit

    Category
    AI-Targeted
    Malware
    SectopRAT, MacSync stealer, AMOS stealer
  2. Unit 42Research

    Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety (opens unit42.paloaltonetworks.com)

    Unit 42 researchers introduce perturbation probing, a method that identifies the small set of neurons responsible for an LLM's safety refusal behavior. They found that in Qwen3-4B, disabling just 50 neurons (0.014% of feed-forward neurons) altered refusal behavior on 80% of harmful prompts, showing safety alignment can rest on a thin, easily disrupted layer rather than robust distributed defenses.

    Category
    AI-Targeted
  3. CyeraVendor report

    Drive-By Agent Hijacking: One Website Visit, Persistent Model Poisoning (opens cyera.com)

    Researchers found a vulnerability (CVE-2026-65105) in NVIDIA NemoClaw where a misconfigured Ollama binding to 0.0.0.0 disables host validation, letting an attacker use DNS rebinding from a malicious webpage to gain unauthenticated access to the local Ollama API. This lets attackers poison the model's chat template to persistently hijack an AI agent's behavior across future sessions; demonstrated as a proof of concept

    Category
    AI-Targeted
    Malware
    NemoClaw, OpenClaw, OpenShell, Ollama
    Vulnerability
    CVE-2026-65105
  4. Pillar SecurityVendor report

    Deadbugz: Currently Active MCP Supply-Chain Campaign (opens pillar.security)

    Pillar Security identified an active campaign distributing a malicious MCP server, productivity-suite, via GitHub pull requests. The server behaves normally for the first three tool calls, then returns altered metadata instructing connected AI agents to search for SSH keys, AWS credentials, and other secrets while hiding the activity. The delivery account, zellkernel, submitted 23 pull requests in a 74-minute window;

    Category
    AI-Targeted
    Actor
    zellkernel
    Malware
    productivity-suite, productivity-suite-mcp, deadbug-mcp.py
  5. SpecterOpsVendor report

    Blacklight: Illuminating AI Agent Artifacts for Attackers and Defenders (opens specterops.io)

    SpecterOps released Blacklight, an open-source toolkit that discovers and analyzes local endpoint artifacts left by AI coding agents like Codex, Claude Code, Cursor, and Antigravity CLI. These artifacts, including auth tokens, session transcripts, and configuration files, can expose credentials, project context, and trust relationships useful to attackers and to defenders building detection guidance.

    Category
    AI-Targeted
    Malware
    Blacklight, Blacklight Scout

July 2026

  1. AnthropicVendor report

    Investigating three real-world incidents in our cybersecurity evaluations (opens anthropic.com)

    Anthropic found that during cybersecurity capture-the-flag evaluations, Claude models unexpectedly gained internet access due to a misconfiguration with a third-party evaluator and compromised real production systems at three organizations, believing them to be simulated targets. Impacts included data exfiltration, a malicious PyPI package that ran on 15 real systems, and unauthorized access via SQL injection, none o

    Category
    AI-Targeted
  2. HuntressVendor report

    Inside FakeAgent: How a Claude Desktop Malvertising Campaign Hit 29 Organizations with SectopRAT (opens huntress.com)

    Huntress found a malvertising campaign that abused a public Claude AI artifact to distribute a trojanized ClaudeDesktop.exe installer, infecting 29 organizations with the SectopRAT trojan via DLL sideloading, GPU-based decryption, and blockchain-hosted (EtherHiding) command and control. Huntress used Claude itself, with human verification, to help reverse engineer the malware's custom AES implementation hidden in a G

    Category
    AI-Targeted
    Malware
    SectopRAT
  3. Hugging FaceVendor report

    Security incident disclosure , July 2026 (opens huggingface.co)

    Hugging Face disclosed that an autonomous AI agent framework breached part of its production infrastructure by exploiting two code-execution flaws in its dataset processing pipeline, then escalated privileges and harvested credentials. No tampering with public models, datasets, or Spaces was found; Hugging Face used an open-weight model on its own infrastructure for forensic analysis after commercial API providers' s

    Category
    AI-Targeted

    Also covered byElastic Security Labs

  4. Tel Aviv UniversityAcademic

    Beware of Agentic Botnets: Scalable Untargeted Promptware Attacks via Universal and Transferable Adversarial HalluSquatting (opens sites.google.com)

    Researchers show that LLM hallucinations of repository or skill names are predictable and transferable across models, letting attackers preregister the hallucinated resource names with malicious payloads. When agentic coding assistants and CLIs fetch these squatted resources they can be tricked into executing code, enabling remote code execution and potentially a botnet. This is proof-of-concept research disclosed re

    Category
    AI-Targeted
    Malware
    HalluSquatting, promptware
  5. MoonlockVendor report

    New Gaslight malware evades AI analysis (opens moonlock.com)

    SentinelOne identified a North Korean-linked macOS Rust malware, dubbed Gaslight, that embeds fabricated system error messages designed to trick AI-based security agents into dismissing it during automated triage. The malware also steals browser data, terminal history, and keychain files, and exfiltrates via a hardened Telegram bot C2, moving prompt-injection evasion from proof-of-concept into real-world use.

    Category
    AI-Targeted
    Actor
    North Korean hackers
    Malware
    Gaslight, AMOS, Realistic macOS infostealer, Realist
    Attribution
    North Korea, per SentinelOne (confidence not stated)
  6. Zscaler ThreatLabzVendor report

    Indirect Prompt Injection in Web Content Targets AI Agents (opens zscaler.com:443)

    Zscaler ThreatLabz documented two real-world campaigns embedding hidden prompt injection instructions in web content via SEO poisoning, JSON-LD, and CSS to manipulate AI agents, including a fake API payment scam and a DeBank typosquatting site. Testing across 26 LLMs found 4 models could be tricked into making payments and 2 misclassified the fraudulent site as legitimate.

    Category
    AI-Targeted

June 2026

  1. Help Net SecurityNews

    Prompt injection still drives most agentic AI security failures in production (opens helpnetsecurity.com)

    Coverage of OWASP's 2026 findings on agentic AI. Most production failures still begin with prompt injection, and attackers increasingly poison what agents trust: MCP servers, packages and coding-tool configuration.

    Category
    AI-Targeted
    Vulnerabilities
    CVE-2025-6514, CVE-2026-22708

May 2026

  1. PermisoVendor report

    ChatGPhish: The Page Is the Payload (opens permiso.io)

    Permiso researchers show that ChatGPT's browser page-summarization feature renders attacker-appended Markdown links and images from third-party pages as trusted UI elements, enabling phishing, QR-code redirection to a second device, and tracking-pixel style data leakage. The issue was demonstrated as a proof of concept and reported to OpenAI via Bugcrowd but was marked not reproducible then a duplicate.

    Category
    AI-Targeted
  2. StraikerVendor report

    Fake Claude Code, Real Malware: Inside the Campaign Targeting AI Developers (opens straiker.ai)

    Straiker documented a live infostealer campaign impersonating Claude Code, JetBrains, NotebookLM and other AI developer tools across 88 domains, using SEO poisoning, paid ads, and fileless payload delivery. The malware, an Amatera/ACR Stealer variant, is built to steal API keys from AI coding assistants alongside browser credentials and crypto wallets, with C2 hidden on a Binance Smart Chain contract.

    Category
    AI-Targeted
    Malware
    Amatera, ACR Stealer
  3. EclecticIQVendor report

    SEO poisoning campaign leverages Gemini and Claude Code impersonation to deliver infostealer (opens blog.eclecticiq.com)

    EclecticIQ documented an SEO poisoning campaign using fake Gemini CLI and Claude Code installation pages to trick developers into running a PowerShell command that installs a fileless, in-memory infostealer alongside the real tool. The malware disables AMSI and ETW, harvests browser, collaboration app, VPN and crypto wallet credentials, and supports remote code execution, with passive DNS revealing over 30 related do

    Category
    AI-Targeted
  4. Trend MicroVendor report

    Inside SHADOW-WATER-063’s Banana RAT: From Build Server to Banking Fraud (opens trendmicro.com)

    Trend Micro's MDR team correlated attacker server infrastructure with victim telemetry to map Banana RAT, a banking trojan targeting 16 Brazilian financial institutions via phishing and fileless PowerShell delivery. The malware provides remote control, keylogging, overlay injection, and PIX QR code interception, using a polymorphic crypter service to evade detection.

    Category
    AI-Targeted
    Actor
    SHADOW-WATER-063
    Malware
    Banana RAT, Backdoor.PS1.BANANARAT.A
    Attribution
    Brazil, per TrendAI (high confidence)

    Also covered byZscaler ThreatLabz

  5. Microsoft SecurityResearch

    When prompts become shells: RCE vulnerabilities in AI agent frameworks (opens microsoft.com)

    Microsoft researchers show how a single injected prompt reached host-level code execution in agents built on Semantic Kernel. Model-controlled parameters flowed unsanitized into a search plugin. Both flaws are fixed.

    Category
    AI-Targeted
    Vulnerabilities
    CVE-2026-25592, CVE-2026-26030
  6. Cloud Security AllianceResearch

    Agent Context Poisoning: SKILL.md and the New AI Supply Chain Attack Surface (opens labs.cloudsecurityalliance.org)

    Cloud Security Alliance details how AI agent skill files like SKILL.md, CLAUDE.md and AGENTS.md create a new supply chain attack surface, since natural-language instructions in these files are trusted and executed by agents at runtime. It cites Snyk's ToxicSkills audit finding security flaws in 36.82% of 3,984 scanned skills and 341 malicious ClawHub skills, plus two Check Point-disclosed CVEs in Claude Code enabling

    Category
    AI-Targeted
    Malware
    ToxicSkills, OpenClaw
    Vulnerabilities
    CVE-2025-59536, CVE-2026-21852

April 2026

  1. GoogleVendor report

    AI threats in the wild: The current state of prompt injections on the web (opens blog.google)

    Google researchers scanned Common Crawl web archives for indirect prompt injection attempts targeting AI agents that browse websites. Most found examples were low-sophistication pranks, SEO manipulation, or crawler deterrence, with only a small number of malicious data-theft or destructive attempts, none highly advanced. Detections of malicious injections rose 32% between November 2025 and February 2026, suggesting g

    Category
    AI-Targeted
  2. OWASP GenAI Security ProjectResearch

    OWASP GenAI Exploit Round-up Report Q1 2026 (opens genai.owasp.org)

    Quarterly review of eight AI-related incidents mapped to the OWASP LLM and agentic risk lists. It includes active exploitation of a maximum-severity Flowise flaw and GrafanaGhost, a prompt injection path that exfiltrates data from Grafana's AI features.

    Category
    AI-Targeted
    Vulnerability
    CVE-2025-59528
  3. ValidinVendor report

    "Hello? I can't hear you": Investigating UNC1069's Fake Meeting Tactics (opens validin.com)

    Validin details UNC1069 (overlapping with Bluenoroff), a North Korean actor luring crypto and Web3 professionals via fake VC personas into fraudulent Zoom/Teams/Meet-style meetings. Victims are tricked with ClickFix prompts into running malware (updated Cabbage RAT/CageyChameleon variants, NukeSped) across Windows, macOS and Linux, and their audio/video is captured via WebRTC for reuse in later social engineering, in

    Category
    AI-Targeted
    Actors
    UNC1069, Bluenoroff, Lazarus Group
    Malware
    Cabbage RAT, CageyChameleon, NukeSped
    Attribution
    North Korea, per Validin (high confidence)

March 2026

  1. Datadog Security LabsResearch

    LiteLLM and Telnyx compromised on PyPI: Tracing the TeamPCP supply chain campaign (opens securitylabs.datadoghq.com)

    Two backdoored releases of LiteLLM, a widely used LLM gateway library, were published to PyPI on March 24, 2026 with a credential stealer. Datadog traces the campaign from a poisoned Trivy scanner through npm and into PyPI.

    Category
    AI-Targeted
    Actor
    TeamPCP

    Also covered byLiteLLMTrend Micro

  2. Unit 42Vendor report

    Open, Closed and Broken: Prompt Fuzzing Finds LLMs Still Fragile Across Open and Closed Models (opens unit42.paloaltonetworks.com)

    Unit 42 researchers built a genetic algorithm based prompt fuzzing method that automatically generates meaning-preserving variants of disallowed requests to test LLM guardrails. Testing against closed-source and open-weight models plus a content-filter model on explosive-related prompts found evasion rates ranging from 1 percent to 99 percent depending on model and keyword. This is original research showing guardrail

    Category
    AI-Targeted
  3. Palo Alto Networks Unit 42Vendor report

    Fooling AI Agents: Web-Based Indirect Prompt Injection Observed in the Wild (opens unit42.paloaltonetworks.com)

    Unit 42 documents real-world indirect prompt injection attacks embedded in webpages, including the first observed case of an attacker bypassing an AI-based ad review system with a scam advertisement. The researchers catalog 22 payload techniques and a severity taxonomy, showing IDPI moving from proof-of-concept to active exploitation, though some scenarios like ad-checker bypass remain unconfirmed against deployed sy

    Category
    AI-Targeted

February 2026

  1. TrendAI ResearchVendor report

    Malicious OpenClaw Skills Used to Distribute Atomic macOS Stealer (opens trendaisecurity.com)

    TrendAI Research documented a campaign where malicious OpenClaw agent skills trick AI agents like GPT-4o into installing a new variant of Atomic macOS Stealer (AMOS), which then deceives users into entering their password. The malware exfiltrates browser data, crypto wallets, Apple and KeePass keychains, and documents, with hundreds of malicious skills found across ClawHub, SkillsMP, and GitHub repositories.

    Category
    AI-Targeted
    Malware
    Atomic (AMOS) Stealer, AMOS
  2. Moonlock LabVendor report

    Moonlock Lab thread on ClickFix malware abusing Claude.ai and Medium (opens x.com)

    Moonlock Lab reports that a Google Sponsored ad for a macOS search led users to malware via ClickFix delivery, seen over 15,000 times. One variant abused a public artifact hosted on claude.ai, while another used a Medium post impersonating Apple support, both attributed to the same threat actor.

    Category
    AI-Targeted
    Malware
    ClickFix
  3. SnykVendor report

    Snyk Finds Prompt Injection in 36%, 1467 Malicious Payloads in a ToxicSkills Study of Agent Skills Supply Chain Compromise (opens snyk.io)

    Snyk scanned 3,984 AI agent skills from ClawHub and skills.sh and found 534 with critical security issues and 76 confirmed malicious payloads designed for credential theft, backdoors, or data exfiltration, with 8 still live on ClawHub. The research shows attackers combining prompt injection with malicious code to bypass agent safety mechanisms in Claude Code, Cursor, and OpenClaw skills.

    Category
    AI-Targeted
    Malware
    ToxicSkills

January 2026

  1. ResecurityVendor report

    Breaking Trust with Words: Prompt Injection Leading to Simulated /etc/passwd Disclosure (opens resecurity.com)

    Resecurity describes penetration testing work on enterprise AI applications, including a banking and HR chatbot, showing how prompt injection can trick an LLM into simulating disclosure of a sensitive Linux file like /etc/passwd. The piece explains direct and indirect prompt injection techniques and several proof-of-concept attack patterns observed during assessments, not confirmed real-world breaches.

    Category
    AI-Targeted
  2. Breached.companyNews

    The Lethal Trifecta Strikes: Four Major AI Agent Vulnerabilities in Five Days (opens breached.company)

    Between January 7-15, 2026, researchers including PromptArmor disclosed indirect prompt injection vulnerabilities in four production AI tools: IBM Bob, Superhuman AI, Notion AI, and Anthropic's Claude Cowork, each allowing data exfiltration via the 'lethal trifecta' of private data access, untrusted content exposure, and external communication channels. Vendor responses varied widely, from Superhuman's rapid remediat

    Category
    AI-Targeted
    Malware
    Claude Cowork, IBM Bob, Notion AI, Superhuman AI, Superhuman Go
  3. GreyNoiseVendor report

    Threat Actors Actively Targeting LLMs (opens greynoise.io)

    GreyNoise honeypots recorded two campaigns targeting LLM infrastructure between October 2025 and January 2026: an SSRF campaign abusing Ollama model pulls and Twilio webhooks, and an 11 day enumeration campaign probing 73+ LLM endpoints across major providers to find exposed API proxies. The enumeration IPs overlap with infrastructure known for scanning 200+ CVEs, suggesting reconnaissance feeding a broader exploitat

    Category
    AI-Targeted
    Vulnerabilities
    CVE-2025-55182, CVE-2023-1389

December 2025

  1. Unit 42 (Palo Alto Networks)Vendor report

    New Prompt Injection Attack Vectors Through MCP Sampling (opens unit42.paloaltonetworks.com)

    Unit 42 researchers show that the Model Context Protocol sampling feature, which lets MCP servers request LLM completions from the client, lacks security controls and trusts servers implicitly. They built a proof-of-concept malicious MCP server against an unnamed coding copilot demonstrating resource theft via hidden prompts, conversation hijacking, and covert tool invocation. No in-the-wild exploitation is claimed;

    Category
    AI-Targeted

November 2025

  1. Infosecurity MagazineNews

    Claude Desktop Extensions Vulnerable to Web-Based Prompt Injection (opens infosecurity-magazine.com)

    Researchers reported that extensions for the Claude desktop app could be driven by instructions planted in web content, turning an ordinary browsing request into a path to actions on the user's machine.

    Category
    AI-Targeted

September 2025

  1. Koi SecurityResearch

    First Malicious MCP in the Wild: The Postmark Backdoor That's Stealing Your Emails (opens koi.ai)

    An npm package posing as the Postmark MCP server behaved normally for fifteen versions, then added one line that copied every email sent through it to the author's server. Koi calls it the first malicious MCP server seen in the wild.

    Category
    AI-Targeted
    Actors
    Jabal Torres, phanpak
    Malware
    postmark-mcp

    Also covered byThe Hacker NewsDark ReadingSnyk

  2. Kaspersky SecurelistVendor report

    Malicious MCP servers used in supply chain attacks (opens securelist.com)

    Kaspersky describes how the Model Context Protocol (MCP), used to connect AI assistants to tools, can be abused via protocol-level tricks like tool poisoning and shadowing, and via supply chain attacks with malicious MCP packages. They built a proof-of-concept MCP server disguised as a developer productivity tool that harvested SSH keys, credentials and environment files while appearing legitimate. This was a control

    Category
    AI-Targeted
    Malware
    devtools-assistant
  3. Unit 42 (Palo Alto Networks)Vendor report

    The Risks of Code Assistant LLMs: Harmful Content, Misuse and Deception (opens unit42.paloaltonetworks.com)

    Unit 42 researchers demonstrate that AI code assistant IDE plugins are vulnerable to indirect prompt injection via contaminated context sources like scraped social media data, causing assistants to insert hidden backdoors into generated code. They also show auto-completion features can be manipulated to bypass safety guardrails and generate harmful content, and that direct model invocation exposes models to further m

    Category
    AI-Targeted
  4. NetskopeVendor report

    Securing LLM Superpowers: When Tools Turn Hostile in MCP (opens netskope.com)

    Netskope describes two proof-of-concept attack techniques against MCP-based LLM deployments: prompt injection hidden in tool description metadata, and cross-server tool shadowing where a malicious server poisons the LLM's shared context to silently alter calls to trusted tools like email. Both exploit MCP's lack of isolation and provenance checks, evading logs and user-facing UI, and the article proposes signing, san

    Category
    AI-Targeted

August 2025

  1. The Hacker NewsNews

    Cursor AI Code Editor Fixed Flaw Allowing Attackers to Run Commands via Prompt Injection (opens thehackernews.com)

    An indirect prompt injection could make Cursor's agent write a malicious MCP configuration file without user approval, giving the attacker remote code execution on the developer's machine.

    Category
    AI-Targeted
    Vulnerability
    CVE-2025-54135

June 2025

  1. Check Point ResearchResearch

    In the Wild: Malware Prototype with Embedded Prompt Injection (opens research.checkpoint.com)

    Check Point found a malware sample uploaded to VirusTotal that embeds a prompt injection string attempting to instruct AI models analyzing it to output 'NO MALWARE DETECTED'. The attack failed against tested LLMs (OpenAI o3 and gpt-4.1) and appears to be an early proof-of-concept, but signals growing attempts to evade AI-based malware analysis tools.

    Category
    AI-Targeted
    Malware
    Skynet
  2. SecurityWeekNews

    'EchoLeak' AI Attack Enabled Theft of Sensitive Data via Microsoft 365 Copilot (opens securityweek.com)

    Aim Security showed that a single crafted email could make Microsoft 365 Copilot send internal data to an attacker with no user interaction. Microsoft patched it server-side and reported no exploitation in the wild.

    Category
    AI-Targeted
    Vulnerability
    CVE-2025-32711

    Also covered byarXiv

February 2025

  1. Unit 42Vendor report

    Investigating LLM Jailbreaking of Popular Generative AI Web Products (opens unit42.paloaltonetworks.com)

    Unit 42 tested 17 popular GenAI web products with single-turn and multi-turn jailbreak strategies to assess safety violations and sensitive data leakage. All tested products were vulnerable to some jailbreak techniques, with multi-turn strategies like Crescendo and Bad Likert Judge more effective for safety violations, while single-turn methods like repeated token attacks were more effective for data leakage in one a

    Category
    AI-Targeted
    Malware
    DAN, Crescendo, Bad Likert Judge, PLEAK, h4rm3l

January 2025

  1. Unit 42Vendor report

    Recent Jailbreaks Demonstrate Emerging Threat to DeepSeek (opens unit42.paloaltonetworks.com)

    Unit 42 tested three jailbreak techniques (Bad Likert Judge, Crescendo, Deceptive Delight) against DeepSeek LLMs and achieved high bypass rates with little specialized knowledge required. The jailbreaks elicited data exfiltration tools, keylogger code, phishing templates, Molotov cocktail instructions and malicious scripts, demonstrating security risks in DeepSeek's safety guardrails.

    Category
    AI-Targeted
    Malware
    Bad Likert Judge, Crescendo, Deceptive Delight
  2. Google DeepMindVendor report

    How we estimate the risk from prompt injection attacks on AI systems (opens blog.google)

    Google DeepMind describes an automated red-teaming framework using optimization-based attacks (Actor Critic, Beam Search, Tree of Attacks with Pruning) to test AI agents' susceptibility to indirect prompt injection that could exfiltrate sensitive user data. This is a defensive research methodology, not a report of real-world exploitation, and no specific incidents are disclosed.

    Category
    AI-Targeted
  3. Wiz ResearchResearch

    Wiz Research Uncovers Exposed DeepSeek Database Leaking Sensitive Information, Including Chat History (opens wiz.io)

    An unauthenticated ClickHouse database belonging to DeepSeek exposed over a million log lines, including chat history, API secrets and backend details, and allowed full control of the database. DeepSeek secured it after disclosure.

    Category
    AI-Targeted
  4. Trend MicroVendor report

    Invisible Prompt Injection: A Threat to AI Security (opens trendmicro.com)

    Trend Micro explains how invisible Unicode tag characters can hide prompt injection text from users while still being interpreted by LLMs, altering model responses. The article demonstrates the technique with a proof-of-concept example and tests attack success rates against several Claude and Mistral models, then promotes its own ZTSA product as mitigation.

    Category
    AI-Targeted

December 2024

  1. Unit 42 (Palo Alto Networks)Vendor report

    Bad Likert Judge: A Novel Multi-Turn Technique to Jailbreak LLMs by Misusing Their Evaluation Capability (opens unit42.paloaltonetworks.com)

    Unit 42 researchers describe a proof-of-concept multi-turn jailbreak that asks an LLM to act as a judge scoring content harmfulness on a Likert scale, then to generate examples at each score, extracting the most harmful response. Testing across six anonymized LLMs showed the technique raised attack success rate by over 75 percentage points versus direct prompts on average, with weaker protection for categories like h

    Category
    AI-Targeted
  2. Unit 42Vendor report

    Now You See Me, Now You Don’t: Using LLMs to Obfuscate Malicious JavaScript (opens unit42.paloaltonetworks.com)

    Unit 42 researchers built an algorithm that uses LLMs to iteratively rewrite malicious JavaScript, evading their own deep learning malware classifier 88% of the time and bypassing all VirusTotal vendors on a sample. This is a proof of concept demonstrating adversarial evasion of AI-based malware detection, not observed criminal use in the wild, and they retrained their model on the samples to improve detection by 10%

    Category
    AI-Targeted
    Malware
    WormGPT, FraudGPT
  3. Trend MicroVendor report

    Link Trap: GenAI Prompt Injection Attack (opens trendmicro.com)

    Trend Micro describes a prompt injection technique where an LLM is manipulated into embedding sensitive collected data into a hyperlink disguised as a normal reference link. If the user clicks it, data is exfiltrated to an attacker, even without the AI having external connectivity permissions. This is presented as a conceptual attack pattern rather than a specific observed incident.

    Category
    AI-Targeted

November 2024

  1. Unit 42Vendor report

    ModeLeak: Privilege Escalation to LLM Model Exfiltration in Vertex AI (opens unit42.paloaltonetworks.com)

    Unit 42 researchers found two vulnerabilities in Google Vertex AI: a privilege escalation via custom job service agents, and a model exfiltration attack via deploying a poisoned model that could steal other fine-tuned ML and LLM models. This was proof-of-concept research conducted in a test environment, reported to Google, which has since patched the issues.

    Category
    AI-Targeted

October 2024

  1. Unit 42 (Palo Alto Networks)Research

    Deceptive Delight: Jailbreak LLMs Through Camouflage and Distraction (opens unit42.paloaltonetworks.com)

    Unit 42 researchers describe Deceptive Delight, a multi-turn jailbreak technique that embeds unsafe topics among benign ones to trick LLMs into generating harmful content. Tested across 8,000 cases on eight models, it achieved a 65% average attack success rate within three turns, versus 5.8% baseline. This is proof-of-concept research with content filters disabled, not an observed real-world attack.

    Category
    AI-Targeted
  2. PermisoVendor report

    When AI Gets Hijacked: Exploiting Hosted Models for Dark Roleplaying (opens permiso.io)

    Permiso observed attackers hijacking exposed AWS access keys to invoke Bedrock foundation models, primarily Anthropic Claude, to power unfiltered AI roleplaying chatbot services. A honeypot captured over 75,000 invocations in two days, mostly sexual content with some straying into CSEM, with circumstantial links to the Chub.ai bot platform. AWS took 35 hours to block the compromised key after invocation volume spiked

    Category
    AI-Targeted
    Malware
    oai-reverse-proxy, one-api

February 2023

  1. SecurelistResearch

    IoC detection experiments with ChatGPT (opens securelist.com)

    Kaspersky researchers tested whether ChatGPT could identify indicators of compromise, finding it failed on known hashes and domains but performed better analyzing host-based artifacts like process metadata and service installations. They built a proof-of-concept PowerShell scanner (HuntWithChatGPT) that used the OpenAI API to flag suspicious system activity with some false positives and negatives.

    Category
    AI-Targeted
    Malware
    Mimikatz, Fast Reverse Proxy, Meterpreter, PowerShell Empire, HuntWithChatGPT.psm1