PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

Pengfei He
Vishesh Sharma
Ash Fox
George Lee
Jiliang Tang
2026

Abstract

Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and dynamic environments, but this also introduces severe security risks. In particular, indirect prompt injection attacks can compromise agents through malicious instructions hidden in external sources such as web pages, emails, and retrieved documents. Existing defenses are largely reactive, while current automated red-teaming methods mainly optimize attack success rather than systematically uncovering hidden vulnerabilities within the agent pipeline. In this work, we propose PI-Hunter, an automated agentic red-teaming framework that shifts the focus from attack optimization to vulnerability exposure. By combining static attack-surface analysis, source-aware seeding, trajectory evaluation, and feedback-guided exploration, PI-Hunter proactively discovers vulnerable ingestion paths and localizes how malicious instructions propagate through agent reasoning. Extensive experiments across multiple benchmarks, agent architectures, attacks, and defenses show that \method~substantially improves vulnerability exposure and attack-surface coverage compared with existing automated red-teaming baselines, while remaining effective even under strong prompt injection defenses.

Research Areas

×