threat_intelligence14407 wordsRead on Arc Codex

Pwnagent: a knowledge-guided multi-agent system for automatic exploit generation

Abstract Automatic Exploit Generation (AEG) plays an important role in proactive assessment of software threats by identifying vulnerabilities and constructing functional payloads. Existing Large Language Model (LLM)-based methods, however, often struggle to reason about complex exploit logic and to perform runtime introspection, leaving a gap between static vulnerability analysis and dynamic memory behavior. We present PwnAgent, an LLM-driven multi-agent framework for end-to-end exploit generation that combines offensive domain knowledge with active runtime introspection. PwnAgent uses a hierarchical knowledge base for multi-stage exploit reasoning and a feedback-driven self-correction engine to calibrate dynamic memory parameters during execution. Because broad Capture The Flag (CTF) benchmarks offer limited binary-exploitation depth and pwn-specific evaluation must balance reproducibility, difficulty progression, and exploit diversity, we construct a 66-task pwn benchmark from public CTF-style challenges. The benchmark is primarily composed of Linux x86/x86-64 ELF binaries and stack-oriented tasks, with smaller format-string, heap, integer-overflow, ARM, and MIPS subsets used as limited probes beyond the dominant setting. Under the same recent Kimi-K2.6 backend, PwnAgent achieves a 62.12% end-to-end success rate, compared with 31.82% for the evaluated PwnGPT baseline, a 30.30 percentage-point gain. These paired results indicate that structured knowledge guidance, execution-grounded measurement, and feedback repair improve LLM-based exploit generation in the evaluated setting, while the absolute success rate shows that fully autonomous exploitation remains challenging. Similar content being viewed by others Introduction Software vulnerabilities remain a fundamental security risk across cyberspace, particularly in mobile and Internet-of-Things (IoT) environments (Klischies et al. 2025). In 2025, disclosed vulnerabilities reached a record 49,000, representing a 20.8% year-over-year increase (Leverett 2025; Kaniewski et al. 2025). Recent Large Language Model (LLM)-agent systems have demonstrated autonomous exploitation capabilities on real-world websites and even zero-day vulnerabilities (Zhu et al. 2024; Fang et al. 2024). These advances make the timely identification and careful validation of exploitable vulnerabilities increasingly important for proactive defense (Peng et al. 2025). Yet manual exploit development is still highly labor-intensive: it requires a deep understanding of program logic and strong expertise in bypassing mitigations, imposing substantial cognitive and time burdens on security analysts (Kurmus et al. 2025). Although Automatic Exploit Generation (AEG) was introduced more than a decade ago, it has yet to see broad practical deployment. Early frameworks (e.g., APEG (Brumley et al. 2008), AEG (Avgerinos et al. 2011)) pioneered program-verification-based generation, but were limited by the poor scalability of symbolic execution. Subsequent systems (Huang et al. 2012; Cha et al. 2012) improved constraint solving for stack-based exploits, but often assumed that security mitigations were disabled. As defenses advanced, more sophisticated techniques (Hu et al. 2015; Bao et al. 2017; Ispoglou et al. 2018) automated code-reuse attacks to bypass Address Space Layout Randomization (ASLR) and Control-Flow Integrity (CFI). Recent work has further targeted complex heap layouts in interpreters (Heelan et al. 2019), kernel exploitation primitives (Wu et al. 2018; Chen et al. 2020), and the transition from crashes to exploits (Wang et al. 2018). However, these traditional methods remain largely template-driven and rigid in their logic. They generalize poorly across architectures and lack the fine-grained semantic reasoning required for multi-stage exploitation. Large Language Models (LLMs) have reshaped vulnerability analysis by making automated semantic reasoning possible. Recent frameworks use LLM-guided mutation to support deeper bug discovery (Meng et al. 2024; Wang et al. 2024; Xia et al. 2024), while autonomous agents can automate penetration-testing plans (Deng et al. 2024; Xu et al. 2024). In domain-specific synthesis, LLMs have been used to generate XSS payloads (Khan 2024; Babaey and Ravindran 2025) and smart contract exploits (Xiao et al. 2025). For binary exploitation, PwnGPT (Peng et al. 2025) provides a benchmark for LLM-driven Capture The Flag (CTF) solving. In this work, we study CTF-style Linux ELF pwn challenges as a controlled and reproducible setting for execution-grounded exploit generation, rather than as direct evidence of large-scale real-world software exploitation. Despite this progress, AI-driven binary exploitation still faces a fundamental gap between logical reasoning and physical execution. Current LLMs generate exploits offline through text and lack the execution-grounded perception needed to reconcile static assumptions with non-deterministic runtime states. Closing this gap requires treating exploitation as an execution-grounded closed-loop system. We further characterize this gap through three challenges (C1–C3). Opaque runtime states (C1) require dynamic introspection, Sense, to calibrate layouts such as compiler-induced padding. Insufficient domain knowledge (C2) calls for a rigorous knowledge model, Model, to support multi-stage exploit composition. Coarse-grained feedback (C3) requires fine-grained diagnostics, Diagnose, to provide actionable error attribution rather than blind trial and error. Addressing these dimensions simultaneously can exceed the cognitive and contextual capacity of a monolithic LLM, often causing reasoning errors and hallucinations (Wu et al. 2024; Qian et al. 2024; Wang et al. 2024). To overcome this limitation, we propose PwnAgent, a decoupled multi-agent framework organized around this execution-grounded loop. The central design choice is to make the exploit-generation state explicit and to couple Vulnerability Primitive (VP) identification, strategy planning, GDB-grounded measurement, exploit synthesis, and diagnosis-driven repair within a single binary-pwn-specific closed loop. PwnAgent delegates the Sense, Model, and Diagnose functions to specialized agents and coordinates them to convert raw vulnerabilities into functional Exploit Primitives (EPs). This division of labor reduces cognitive load by reflecting expert collaboration in practice. By combining static modeling with debug-driven dynamic introspection, PwnAgent calibrates VPs to actual memory conditions. It also draws on structured knowledge bases and iterative feedback to support reliable exploit generation while reducing the need for manual intervention. Evaluating these autonomous capabilities requires test environments that can differentiate levels of performance, but existing benchmarks expose a trade-off between breadth and binary-exploitation depth. General-purpose CTF datasets (e.g., InterCode-CTF (Yang et al. 2023), NYU-CTF (Shao et al. 2024)) are valuable because they cover multiple security categories, but pwn tasks form only one part of their broader scope. Binary-focused evaluations such as PwnGPT (Peng et al. 2025) provide a closer reference point for AEG, yet the available pwn task space still tends to contain either introductory stack exercises or highly specialized heap and mitigation-bypass challenges. This makes it difficult to observe gradual changes in the ability of an autonomous exploit-generation system. To support a more focused evaluation, we construct a difficulty-graded benchmark of 66 Linux ELF pwn challenges. The benchmark is explicitly capability-oriented rather than statistically representative of all CTF pwn tasks. It is therefore intentionally stack-heavy: stack overflows are a common entry point for binary exploitation, allow controlled variation of mitigations and exploit methods, and expose the offset-calibration and interaction-synchronization failures targeted by PwnAgent. We include format-string, integer-overflow, and heap cases to broaden the exploit surface, but heap challenges are fewer because they are concentrated in the hard tier and often depend on allocator-version details, menu-state management, and challenge-specific heap shaping. Similarly, the benchmark is primarily x86-based, with a small number of ARM and MIPS binaries used as limited cross-architecture probes rather than as evidence of broad architecture coverage. This design gives a smoother difficulty gradient for analyzing the proposed framework, while its imbalance is an explicit limitation discussed in Section Limitations. We summarize the primary contributions of this paper as follows: - We design and implement PwnAgent, an end-to-end AEG system that executes the full exploit-generation workflow from binary-only program analysis to exploit-script synthesis and local validation. Given only the target binary and optional challenge-provided libc, PwnAgent coordinates static program understanding, exploit strategy planning, dynamic memory measurement, exploit synthesis, and feedback-driven repair. Within this architecture, we integrate a structured pwn knowledge base through role-specific cognitive sharding. The knowledge base provides agents with guidance derived from audited exploitation principles, reasoning rule chains, reference trajectories, and empirical pitfalls, while excluding benchmark-specific write-ups, reference exploits, and challenge artifacts from the deployed prompts. - We build a capability-oriented benchmark of 66 CTF pwn challenges and define a strict evaluation protocol. The benchmark is primarily composed of Linux x86/x86-64 ELF binaries, with limited ARM and MIPS cases used only as cross-architecture probes, and the evaluation counts end-to-end success only when the generated exploit script obtains a shell or reads the flag in the local environment. - We provide an empirical evaluation of PwnAgent against the evaluated PwnGPT baseline and ablations of the knowledge base and dynamic-analysis components. Under the same recent Kimi-K2.6 backend, PwnAgent achieves a 62.12% end-to-end success rate on this benchmark, compared with 31.82% for PwnGPT, corresponding to a 30.30 percentage-point gain; ablation results indicate that structured knowledge guidance and execution-grounded measurement both contribute to the observed gains, while the remaining failures highlight the difficulty of fully autonomous exploitation. Background and motivation Technical preliminaries To establish the terminology used by PwnAgent, we introduce the exploit-generation concepts and agent abstractions that are needed in the subsequent system design. AEG treats exploitation as a series of state transitions within the execution flow of a target program. Following established formalisms in offensive security (Garmany et al. 2018; Liu et al. 2022), we define this process using two core building blocks. The first is the VP, an abstract root-cause capability derived from a software flaw and identifiable via static code semantics (e.g., an out-of-bounds array index or a use-after-free condition). As Liu et al. (2022) note, a VP represents the potential for exploitation but lacks physical grounding or direct control over resources. The second building block is the EP, a functional capability achieved by manipulating the physical memory state. Garmany et al. (2018) define this as crafting specific data objects or payloads to gain stabilized control over critical system invariants (e.g., arbitrary read/write, or hijacking the instruction pointer). Successful autonomous exploitation requires identifying a valid VP and iteratively synthesizing the corresponding EP. This transition must satisfy strict physical invariants imposed by modern defense mechanisms (e.g., Position Independent Executables (PIE), ASLR, Stack Canaries) and architectural alignment conventions. Recent work has made it possible for LLMs to operate as autonomous agents that can make sequential decisions. The ReAct (Reasoning and Acting) framework proposed by Yao et al. (2023) combines an agent’s internal reasoning trace with task-specific external actions, that is, tool use. To support complex logical deduction, such as planning a multi-stage exploit, ReAct relies on Chain-of-Thought (CoT) prompting (Wei et al. 2022), which encourages step-by-step decomposition of the problem to reduce logical errors. To reduce hallucination and incorporate domain knowledge, prior work constructs structured knowledge bases that formally encode offensive security priors (Shao et al. 2025). Although LLMs can handle substantial input, they still face strict context limits. Putting large amounts of domain knowledge, such as dense exploitation techniques, into a single prompt reduces performance. As Liu et al. show for the lost-in-the-middle effect (Liu et al. 2024), LLMs struggle to use relevant information buried in long contexts and tend to favor information near the beginning or end. Retrieval-Augmented Generation (RAG) remains an effective mechanism for selectively injecting external knowledge, but recent studies also show that retrieved contexts can introduce noise, hard negatives, and quality-latency trade-offs (Chen et al. 2024; Jin et al. 2024; Yu et al. 2024; Ray et al. 2025). To manage these limitations in long-horizon agent workflows, recent work has increasingly turned to multi-agent role specialization (Wu et al. 2024; Qian et al. 2024). These approaches decompose complex tasks and distribute them across specialized sub-agents. Under Cognitive Sharding, fragments of the knowledge base are statically embedded in the initialization prompts (i.e., personas) assigned to these agents. Each sub-agent operates within the tactical scope of its role, helping maintain high reasoning density without exceeding context limits. Since zero-shot inference often fails in opaque environments, paradigms such as Reflexion (Shinn et al. 2023) show that language agents can improve performance through explicit closed-loop feedback and iterative self-tuning. In automated programming settings, self-debugging architectures (Chen et al. 2024) allow LLMs to analyze execution results returned by the environment (e.g., process crashes) and systematically refine generated code without human intervention. Motivation: the logical-physical gap Recent advances in LLMs have created new opportunities to address the semantic gap in traditional AEG by improving code understanding and strategic planning. However, applying monolithic, text-only LLMs (e.g., PwnGPT) directly to binary exploitation tasks exposes a fundamental limitation: a serious misalignment between the static semantic plane and the dynamic physical plane. Such models operate as offline text processors over static code representations (the semantic plane) and cannot observe dynamic, non-deterministic execution environments (the physical plane). This mismatch gives rise to the Logical-Physical Gap between high-level strategic reasoning and the low-level physical realities of target binary execution. Bridging this gap and building a reliable end-to-end AEG system requires our architecture to address three central challenges. The first challenge is the introspection deficit (C1). Without active runtime probing, monolithic LLMs must infer memory layouts from static representations and tool outputs. Modern compiler optimizations and system-level protection mechanisms introduce physical invariants into memory, such as stack alignment padding and randomized metadata, that can be difficult to recover reliably from decompiled views alone. As a result, deductions based only on uncalibrated static semantics can fail. To illustrate this limitation, we consider stack2, a typical vulnerability challenge that requires an out-of-bounds (OOB) write to bypass a Stack Canary. As shown in Fig. 1(a), the vulnerability in stack2 is an OOB array update in the change number routine. This update allows an attacker to overwrite stack data byte-by-byte by supplying an unchecked index. Static analysis from decompiled pseudocode locates the array base at ebp-0x70, establishing a naive overwrite index of 0x74 (116) for the target return address based on the standard ebp+0x04 assumption. A planner that relies on this decompiled stack view without verifying the epilogue-induced restoration path can therefore select 0x74 as the overwrite index. The full assembly, however, reveals compiler-introduced prologue and epilogue adjustments (Fig. 1b). To enforce strict stack alignment, the compiler saves the original stack pointer to a register (e.g., ecx) before applying the and esp, 0xfffffff0 instruction. Crucially, during the epilogue, a lea esp, [ecx-4] instruction restores this unmodified stack frame immediately before the ret instruction. This restoration path does not imply that the epilogue is inherently beyond static reasoning; rather, it makes the decompiler-derived ebp+0x04 offset an unsafe proxy for the physical overwrite index. At runtime, the actual return address is fetched from the restored, unaligned high-memory address rather than the local ebp+0x04 boundary, shifting the required overwrite index to 0x84 (132). Thus, even when static evidence for the mechanism is present, the final overwrite index remains a physical exploit parameter that should be calibrated in the execution environment. If the system commits to 0x74 without such calibration, it fails to hijack control flow and may only observe anomalous behavior or an uninformative SIGSEGV crash. Equipping agents with dynamic perception allows them to measure the physical layout directly and reduces payload failures caused by compiler-induced stack restoration and alignment effects. The second challenge is the knowledge representation gap (C2). Exploit development goes beyond conventional programming and requires careful adherence to implicit low-level system rules, including architecture-specific Application Binary Interface (ABI) calling conventions and mitigation bypass techniques (e.g., ASLR, CFI). Although existing monolithic text models are pre-trained on large code corpora, they do not capture this kind of structured offensive and defensive expertise. When faced with complex vulnerability chains that require multi-stage exploitation, these models often hallucinate during reasoning. They may generate attack payloads that appear semantically plausible but still violate low-level physical constraints. Exploring a large state space without prior guidance leads to substantial trial-and-error cost. AEG systems therefore need to incorporate a structured expert knowledge base. By abstracting discrete exploitation experience into machine-executable CoT reasoning paths, such a representation can guide the generation process more reliably. It narrows the decision search space, reduces planning errors, and suppresses hallucinations, thereby improving the reliability of exploit code generation. The third challenge is feedback ambiguity (C3). Execution failures occur frequently during exploit generation. Modern operating systems typically return simplified and uninformative signals, such as process crashes or timeouts, when illegal memory accesses occur. Monolithic agents limited to a static view cannot observe the underlying physical errors when they receive these generic crash signals. As a result, they cannot map physical execution symptoms back to their logical root causes, which leads to incorrect diagnoses, such as blindly adjusting shellcode formats, and can trap the process in unproductive retry loops. Code-generation ability alone is therefore insufficient; models also need a self-debugging and reflection mechanism driven by dynamic feedback. This architecture requires a dedicated diagnostic module that can extract the physical context at the moment of a crash (e.g., register states, memory dumps) and guide subsequent correction. Translating low-level execution anomalies into precise, actionable error attribution reports allows the model to correct parameter deviations or logical flaws and thereby achieve iterative self-repair. Ultimately, bridging the Logical-Physical Gap caused by static blind spots, missing domain knowledge, and feedback ambiguity requires an architecture that combines static model reasoning with runtime dynamic introspection. This architecture must operate under the strict constraints of an expert knowledge base. These requirements form the foundation for proposing the PwnAgent multi-agent collaborative framework. Toward a multi-agent solution Recent advances in LLM research have opened up new directions for AEG. To address the core challenges (C1 to C3), we propose three methodological principles for building an exploit generation architecture with execution awareness and strategic robustness. The first principle centers on environment perception and synergy across multiple tools. Giving LLMs external tool-calling capabilities extends their ability to analyze and verify vulnerabilities. This design separates high-level logical reasoning from environment observation through interactions among agents, security tools, and execution environments. Agents can then autonomously detect program states and obtain context in real time (Deng et al. 2024; Fang et al. 2024). Feedback from the toolchain helps overcome the cognitive limits of text-only models by creating an interactive analysis loop. This loop adapts to program states dynamically and provides accurate execution context for subsequent reasoning. The second principle combines domain priors with CoT guidance. Binary exploit development requires substantial reasoning and careful attention to implicit execution constraints. Encoding discrete exploitation experience in structured knowledge bases improves the reliability of model reasoning (Shao et al. 2025; Yang et al. 2023). To reduce the logical errors common in direct code generation, CoT prompting makes intermediate reasoning steps explicit and helps maintain precision in complex tasks (Wei et al. 2022; Meng et al. 2024). Combined with predefined reasoning templates, CoT decomposition provides stable guidance for primitive identification, technique selection, and staged planning. This helps ground the model’s reasoning in established technical specifications and improves the controllability and completeness of multi-stage exploit strategies. The third principle concerns layout-driven closed-loop feedback and reflection. A robust AEG system requires closed-loop self-correction to address biases that arise during exploitation. To overcome the limitations of coarse-grained feedback signals, this approach builds a reflection loop around layout-oriented feedback (Shinn et al. 2023; Chen et al. 2024). The loop maps anomalies in system state directly to higher-level causes of failure in exploit logic. Within this framework, a generic execution failure becomes diagnostic evidence with precise layout guidance. This enables continued refinement of generation strategies in binary environments with complex mitigations. System design This section details the design philosophy and system architecture of PwnAgent. To address the fundamental challenges of missing runtime information (C1), heavy dependence on domain knowledge (C2), and sparse feedback signals (C3), PwnAgent introduces a multi-agent framework designed to orchestrate the transformation of raw vulnerabilities into functional EPs. The system conceptualizes exploit generation as a closed-loop process comprising binary semantic and interaction analysis to identify VPs, knowledge-driven exploit strategy planning, targeted dynamic memory probing, environment-grounded exploit synthesis, and execution-driven verification and remediation. Problem scoping and threat model To clarify the scope of our study and define the rigorous boundaries under which PwnAgent operates, we explicitly specify the target environment and agent capabilities. This formulation serves as the prerequisite threat model for our autonomous exploitation framework: - Target scope (\(\mathcal {B}\)): We focus on user-space memory corruption vulnerabilities (e.g., stack/heap overflows and format strings) residing within Linux Executable and Linkable Format (ELF) binaries. Kernel-level exploits and browser engine vulnerabilities are considered orthogonal research tracks and lie outside the current scope. - Execution environment: The target binary runs as a Linux ELF program under standard user-space ABI conventions. Our current evaluation is dominated by x86/x86-64 binaries and includes only a small number of ARM and MIPS cases as limited non-x86 probes, rather than as evidence of broad architecture coverage. The system is evaluated in environments protected by widely deployed defense mechanisms, including Non-Executable (NX) memory, Stack Canaries, PIE, and ASLR. - Agent capabilities (Gray-box Setting): We assume the agent possesses local execution privileges, enabling it to iteratively launch the binary via standard I/O streams and to attach dynamic debuggers (i.e., GDB) for runtime introspection. However, we strictly operate in a gray-box setting: the agent does not presume access to source code or proprietary debug symbols beyond the binary provided. - Input assumption: At startup, PwnAgent receives only the target binary and, when the challenge distributes a specific runtime library, the corresponding libc. It is not given the challenge description, source code, debug symbols, flaw-level hints, vulnerability locations, proof-of-concept inputs, crash-triggering inputs, reference exploits, or write-ups. Given this scoped setting, the fundamental objective of PwnAgent is to ingest an initial state tuple \(\Omega _{\textit{init}} = \langle \mathcal {B}, \mathcal {D}_{\textit{lib}}, \mathcal {W}_{\textit{env}} \rangle\)—comprising the target binary \(\mathcal {B}\), optional challenge-provided dynamic libraries \(\mathcal {D}_{\textit{lib}}\) (e.g., a specific libc), and the non-semantic execution workspace description \(\mathcal {W}_{\textit{env}}\) (e.g., file paths and runtime directory layout)—and autonomously generate a verified, functional exploit script \(\mathcal {E}\) (e.g., a pwntools Python script) that reliably obtains an interactive shell or reads the flag in the local evaluation environment. No vulnerability-level information is included in \(\Omega _{\textit{init}}\). Architecture and collaboration paradigm PwnAgent is an LLM-based automatic exploit generation framework that structures the exploitation process through a tiered architecture. As shown in Fig. 2, the system is organized into two functional tiers: the Reasoning Layer and the Execution Layer. This design separates high-level strategic reasoning from low-level, non-deterministic interactions with the environment. By adopting this modular structure, the system keeps strategic planning insulated from the variability of raw tool outputs and maintains consistency throughout the multi-stage exploit generation process. System architecture Unlike traditional multi-agent systems, PwnAgent adopts a Centralized Reasoning, Peripheral Interpretation design paradigm implemented through a ReAct (Yao et al. 2023) feedback loop. The Reasoning Layer is centered on the MainOrchestrator, which serves as the system’s central decision engine. It manages the global system state and uses a distributed hierarchical knowledge base (KB) organized as a series of CoT reasoning frameworks. Rather than functioning as a static template repository, the KB provides structured guidance for the orchestrator’s reasoning process. Under this design, the orchestrator operates in a continuous loop of reasoning and targeted interaction with the environment. Supporting this process, the Execution Layer consists of three specialized sub-agents, ProgramAnalyzer, MeasurementExpert, and ExploitCrafter, each assigned a distinct set of technical tasks by the MainOrchestrator. These agents are not independent decision-makers; instead, they act as stateless executors that carry out domain-specific operations within restricted scopes and provide accurate observations about the environment. ProgramAnalyzer handles semantic extraction and vulnerability localization, MeasurementExpert performs precise parameter measurement (Stage 3) and diagnoses execution failures (Stage 5), and ExploitCrafter translates strategies into functional code. To interact with the Dynamic Environment, the sub-agents use a unified Environment Interaction Bus. This shared interface integrates the Model Context Protocol (MCP) and Command Line Interface (CLI) tools, providing standardized access to the local filesystem, terminal, GDB, and IDA Pro. The bus provides a synchronous channel for exchanging Tool Commands (Action) and state observations (Observation), which are then incorporated into the orchestrator’s reasoning (Thought). This tiered design addresses the information gap between static analysis and runtime reality. The specialized agents focus on extracting signals, while the orchestrator preserves overall logical coherence. Inter-agent communication protocol Communication between the two layers follows a hierarchical delegation pattern. The system adopts a star-shaped communication topology centered on the MainOrchestrator, which is the sole initiator of tasks. Delegation is strictly unidirectional: the orchestrator sends mandates to sub-agents through specific prompts, such as an API Request (Goal, Context) to ProgramAnalyzer or a Task Dispatch (Context) to MeasurementExpert, while the sub-agents remain mutually isolated and cannot communicate with one another or invoke each other recursively. In this interface, the target sub-agent receives an instr parameter that specifies the technical objectives and output constraints, together with an optional ctx parameter that provides execution context, such as binary paths, to the stateless interpreter. This decoupled topology prevents recursive reasoning loops and ensures that implementation-level complexity does not propagate across specialist boundaries. The protocol uses a strict request-response lifecycle to maintain system stability. To reduce contextual overhead, sub-agents operate in isolated sub-sessions, where intermediate reasoning and raw tool outputs remain hidden, and only structured JSON payloads are returned to the orchestrator. This keeps the primary reasoning layer at a high signal-to-noise ratio by separating it from implementation-level noise. Each delegation is also an independent, stateless call. Sub-agents do not retain internal state across tasks, which makes the system more predictable in complex exploitation scenarios. Data exchange between the orchestrator and the agents relies on JSON-serialized payloads with predefined schemas; for example, vulnerability structures are extracted directly into discrete type, location, and primitives fields rather than parsed from unstructured narrative text. Replacing verbose natural-language logs with structured payloads increases information density and helps mitigate LLM context-window limitations in long-horizon planning. After receiving a payload, the orchestrator performs a state-mapping operation in which JSON fields are synchronized directly with the corresponding slots in the Global System State, for example by mapping the vulnerability field to the analyzed_vuln_site property, thereby ensuring consistency across subsequent stages. Operational workflow Built on this architecture, PwnAgent operates as a closed-loop refinement process with five sequential stages. Throughout this process, the MainOrchestrator interacts continuously with the centralized Global System State, systematically writing outputs such as semantic context, attack strategy, and memory layout, while reading the required inputs to preserve temporal consistency across stages. The workflow begins with Semantic Modeling (Stage 1), which extracts security properties and identifies VPs. These outputs then support Exploit Strategy Selection (Stage 2), where the system maps the identified VPs to viable exploit chains through task decomposition. To ground tactical reasoning in physical memory, PwnAgent next performs Dynamic Memory Probing (Stage 3), resolving non-deterministic parameters with byte-level precision. These grounded measurements are then used as hard constraints in Exploit Code Generation (Stage 4). Finally, the system ensures reliability through Stratified Feedback and Remediation (Stage 5), where execution failures trigger diagnosis-driven backtracking to iteratively correct the script until it obtains a shell, reads the flag, or reaches the attempt budget. Hierarchical knowledge base and cognitive sharding To bridge the gap between general model capabilities and specialized exploitation requirements, we construct a structured knowledge base \(\mathcal {K}\) that encodes domain-expert experience into machine-actionable reasoning modules. Knowledge acquisition and transformation The construction of \(\mathcal {K}\) follows a two-phase process: large-scale material curation, followed by LLM-driven structured distillation. Material Collection. We systematically curated a source corpus of 384 high-quality technical documents from the open-source exploit community, including 246 public CTF pwn write-ups from historical non-benchmark challenges, 30 CVE-related advisories and exploit analyses, and 108 official or community tutorials. The corpus covers four major vulnerability families commonly encountered in CTF-style binary exploitation: stack-based buffer overflows, format-string vulnerabilities, heap-related vulnerabilities, and integer overflows. The heap-related portion includes allocator-oriented patterns such as heap overflow, off-by-one corruption, use-after-free, double-free, and tcache/fastbin manipulation. To prevent benchmark-to-KB leakage, we applied a benchmark-overlap exclusion procedure before finalizing the deployed KB. The exclusion list was built from benchmark challenge names, platform identifiers, file names, aliases, and other available task identifiers. During material curation and KB auditing, we removed any material that directly corresponded to an evaluation task, including public write-ups, proof-of-concept inputs, reference exploits, solution scripts, flags, and challenge-specific analysis notes. Ambiguous matches were conservatively excluded. Thus, the deployed KB retains generalized exploitation knowledge while excluding task-specific solutions for the evaluated benchmark instances. The deployed KB used in the evaluation was frozen before running the benchmark, and no benchmark-specific write-up, reference exploit, intended exploit method, or post-hoc rule derived from test outcomes was added during evaluation. Structured Knowledge Extraction. Rather than relying on rigid manual transcription, we developed an extraction engine where a long-context LLM, guided by schema definitions and few-shot demonstrations, systematically processes the raw materials. The core transformation principle is converting narrative execution steps into guided reasoning chains (CoT inquiries). Through a unified extraction pass, the engine processes the texts across four complementary dimensions. It identifies core calling conventions (e.g., x86_64 System V ABI), atomic EPs (e.g., PC control, arbitrary read/write), and protection mechanism invariants to form the Foundational Principles (\(L_1\)). It then performs Reasoning Path Generalization using isomorphism-based deduplication to merge homologous tactical scenarios and separate distinct ones, converting them into parameterized problem chains. This constructs the Tactical Inference Chains and the Meta-Cognitive Root Chain \(\textit{RC}_0\) (\(L_2\)). Additionally, the engine extracts selected end-to-end CoT reasoning trajectories for few-shot learning references to create the Reference Trajectories (\(L_3\)). It also identifies common technical errors and the underlying root causes of these failures to construct self-reflection hooks and operational constraints, which form the Empirical Pitfalls (\(L_4\)). The raw source documents are not directly injected into prompts or retrieved at runtime; only the audited and distilled knowledge items in this four-layer ontology are used by the agents. These CoT-style chains and trajectories are treated as external guidance schemas for organizing exploit knowledge and constraining agent behavior, rather than as verified transcripts of the model’s internal reasoning process. This iterative process yields a stable LLM-synthesized, four-tier ontology (\(\mathcal {K}\)). Finally, via a Human Expert Audit, a security professional executes a manual curation and verification pass over the generated \(L_1-L_4\) grid to check the factual accuracy of the derived domain axioms and reasoning constraints, finalizing the knowledge base for deployment. Knowledge ontology Definition 1 (Domain Knowledge Base) The knowledge base \(\mathcal {K}\) is defined as a 4-tuple: where each \(L_i\) denotes a distinct layer of cognitive abstraction: foundational principles (\(L_1\)), reasoning rules (\(L_2\)), reference trajectories (\(L_3\)), and empirical pitfalls (\(L_4\)). These four layers form a cognitive progression: \(L_1\) establishes the conceptual baseline, \(L_2\) builds reasoning paths on top of \(L_1\), \(L_3\) illustrates how \(L_2\) is applied, and \(L_4\) highlights common errors (pitfalls) together with root-cause-oriented corrective guidance. A concise overview of this four-tier ontology, coupled with concrete examples extracted from the operational knowledge base, is illustrated in Table 1. Foundational Principle Layer (\(L_1\)). \(L_1\) is defined as a set of fundamental security concepts \(\{c_1, c_2, \ldots , c_m\}\). Each concept \(c_i = \langle \textit{name}, \textit{def}, \textit{mech}, \textit{impact} \rangle\) provides a multi-dimensional description of architectural invariants, covering EPs, protection invariants (e.g., NX, PIE, Canary), and ABI calling conventions. Definition 2 (Reasoning Rule Chain) Let \(\mathcal {T} = \{0, 1, \dots , n\}\) be the index set of tactical paradigms, where \(\textit{RC}_0\) denotes the meta-cognitive root chain. A rule chain \(\textit{RC}_i \in L_2\) for \(i \in \mathcal {T}\) is formally defined as a 4-tuple: where \(\Phi _i\) denotes the activation trigger, \(\mathcal {Q}_i = \langle q_1, q_2, \ldots , q_k \rangle\) is an ordered CoT question sequence used to guide tactical derivation, such as checking requisite preconditions and whether they are currently satisfied in memory, \(\mathcal {H}_i\) captures implicit heuristics and strategic preferences (e.g., favoring simpler mitigation bypasses), and \(\mathcal {C}_i\) defines strict redlines for dynamic verification (e.g., prohibiting reliance on offsets derived solely from static analysis). The inclusion of \(\mathcal {H}_i\) and \(\mathcal {C}_i\) ensures that reflection remains constrained rather than devolving into arbitrary decision-making. In the current implementation, \(L_2\) includes a meta-cognitive root chain (\(\textit{RC}_0\)) and 32 distinct core rule chains covering strategy planning, technical inference, measurement methodology, and exploit synthesis. \(L_3\) contains a dozen non-duplicative reference trajectories, while \(L_4\) contains 20 technical pitfalls and 3 general debugging-thinking traps, supplemented by 18 meta-cognitive hooks for reflection and self-checking. These chains are anchored in the meta-cognitive formulation (\(\textit{RC}_0\)) to enforce system-level reasoning. Outside the frozen evaluation setting, the ontology can be extended as the knowledge corpus grows to incorporate new structural paradigms. Reference Trajectory Layer (\(L_3\)). Each reference trajectory in \(L_3\) records three fields: the exploitation scenario, an audited demonstration trace, and the reusable tactical pattern abstracted from that trace. These entries provide few-shot guidance for strategy composition without exposing benchmark-specific solutions. Empirical Pitfall Layer (\(L_4\)). \(L_4\) represents the set of empirical pitfalls, denoted as \(\mathcal {P}\). Each entry records an observable anomaly, its root cause, and a prevention or remediation rule, such as mapping a 64-bit movaps crash to stack misalignment and the insertion of an alignment ret gadget. In terms of technique coverage, the finalized KB covers representative exploit techniques including ret2text, ret2shellcode, ret2libc, ROP/SROP, ret2csu, stack pivoting, canary leakage and bypass, PIE/libc leakage, GOT/PLT hijacking, format-string arbitrary read/write, one-gadget constraint handling, ORW chains, and heap-oriented techniques such as tcache poisoning, fastbin/tcache double-free variants, unsorted-bin leak/attack, unsafe unlink, House of Force, and hook overwrite. It also encodes architecture-specific calling-convention rules primarily for x86/x86-64, with limited ARM/MIPS syscall and shellcode notes. The full ontology \(\mathcal {K}\) is stored as a specialized Markdown document with explicit delimiters. This plain-text format aligns naturally with the long context windows of modern LLMs, enabling deterministic static sharding without a per-query retrieval step. Cognitive knowledge sharding The knowledge base \(\mathcal {K}\) is distributed under a Cognitive Sharding model, in which specialized knowledge fragments are statically injected into each agent’s system prompt during initialization. We do not claim that this static design is generally superior to RAG. Instead, it is adopted because PwnAgent uses a fixed, audited, and role-specific KB within a deterministic five-stage exploitation workflow. In this setting, the recurring knowledge needs of each role are known before execution: ProgramAnalyzer repeatedly needs program-analysis and ABI constraints, MeasurementExpert needs probing and diagnostic rules, ExploitCrafter needs implementation constraints, and MainOrchestrator needs strategic rule chains and reference trajectories. Static role-specific sharding therefore provides deterministic access to mandatory knowledge, reduces retrieval uncertainty, and lowers the risk that irrelevant or conflicting technique fragments enter a role’s local context. Avoiding runtime query generation, retrieval, ranking, and context assembly also reduces per-iteration overhead, but this is a secondary engineering benefit rather than the primary motivation. For implementation, Cognitive Sharding is specified by a role-to-shard assignment \(\mathcal {S}(r)\), where r denotes an agent role and \(\mathcal {S}(r)\) denotes the fixed subset of \(\mathcal {K}\) injected into that role’s system prompt. During human expert audit, each knowledge item is tagged with the roles for which it is mandatory or directly relevant. The deployed assignment follows two rules: each role receives the concepts, rule chains, trajectories, and pitfalls needed for its delegated tasks, and knowledge outside the role’s responsibility is excluded unless the MainOrchestrator passes it later through schema-constrained state. Table 2 gives the concrete mapping used in the evaluation. By explicitly binding each role to predefined knowledge boundaries through static delimiter-based indexing, such as explicit line-range mappings, the system reduces contextual load. The deployed ontology contains approximately 68.9K estimated tokens before role-specific slicing, and each agent receives only the shard required by its responsibility. Because knowledge items relevant to multiple responsibilities are intentionally assigned to more than one role, the role-specific shards are not mutually exclusive; therefore, the sum of the per-role estimates in Table 2 can exceed the size of the full ontology before slicing. In the implementation, MainOrchestrator is loaded with strategy-planning rules and reference trajectories ( \(\sim\) 33.3K tokens), whereas MeasurementExpert receives only measurement, debugging, and diagnostic knowledge ( \(\sim\) 5.4K tokens). This decoupled design reduces context noise and helps mitigate the lost-in-the-middle context decay commonly observed in LLMs. This cognitive sharding corresponds to the procedural responsibilities of the agents, allowing them to function as domain experts without unnecessary context expansion. By separating abstract strategy frameworks (\(\textit{RC}_0\)) from concrete diagnostic redlines (\(\mathcal {P}\)) at the structural level, the system enforces disciplined CoT reasoning. For example, the verification protocols (\(\mathcal {C}_i\)) assigned to MeasurementExpert focus exclusively on empirical validation of dynamic states, ensuring that the agent’s reasoning remains tightly grounded in the specified measurement directives rather than drifting toward unsupported exploit strategies. As summarized in Table 2, this explicit mapping from static knowledge shards to their corresponding agent roles tightly bounds cognitive scope and reduces the token footprint of each LLM session. This design also has an important trade-off. Since sub-agents are mutually isolated, a role-specific shard can omit contextual facts that become relevant later in the exploit trajectory. PwnAgent mitigates this risk by centralizing cross-stage memory in the MainOrchestrator rather than relying on direct sub-agent communication. The orchestrator stores ProgramInfo, Strategy, Measurements, Diagnostics, and failure history in the Global System State, then forwards only the necessary facts through schema-constrained JSON payloads and the ctx field of each delegated call. Thus, Cognitive Sharding is used as a controlled context-governance mechanism: it narrows local reasoning scope while preserving global temporal consistency through the orchestrator. A hybrid design that combines static role-specific shards with retrieval for rare or out-of-shard cases remains a promising extension, but is outside the scope of the current evaluation. Synergistic generation workflow To solve complex pwn challenges, PwnAgent adopts the Centralized Reasoning, Peripheral Interpretation architecture described in Section Architecture and Collaboration Paradigm. This architecture coordinates an ongoing cycle between the centralized reasoning engine and the specialized peripheral interpreters. The generation workflow consists of five sequential stages, progressing from raw program analysis to exploit code synthesis. This modular design helps ground each stage of the attack in expert knowledge and empirical evidence, while preserving strict data-flow boundaries and clearly defined agent responsibilities. The five stages are described below. Stage 1: Programanalyzer-driven semantic modeling In Stage 1, the system performs a comprehensive static analysis of the target binary \(\mathcal {B}\) to localize candidate vulnerabilities from the binary alone and evaluate their potential impact. The process begins when the orchestrator assigns a static analysis task (①) to ProgramAnalyzer, a stateless static inference engine that is fully isolated from dynamic execution and debugging environments. Using IDA Pro through the MCP interface, together with command-line tools such as checksec, the agent examines the decompiled pseudocode and binary metadata from multiple perspectives without relying on a supplied proof-of-concept input, crash trigger, or vulnerability location. Rather than functioning as a rudimentary pattern-matching preprocessor, the agent identifies three critical aspects. First, through Vulnerability Capability Deduction, it deduces the operational capabilities of a discovered flaw (e.g., translating a subverted format string into an arbitrary read/write primitive). Second, via Interaction Flow Modeling, it statically traces control paths and extracts contiguous string literals to model the deterministic menu prompts and input triggers to reach the vulnerable state. Third, through Constraint Identification, it extracts local structural limitations, such as buffer capacities, forbidden byte filters, and baseline binary mitigations (e.g., NX, Canary). The agent structures these static findings into the high-fidelity semantic model \(\mathcal {I}\). It then passes this comprehensive understanding of the nature of the vulnerability, the requirements of the state machine, and theoretical limitations to the central orchestrator as structured semantic results (②) for subsequent strategy formulation. Stage 2: Orchestrator-driven exploit strategy selection Based on the semantic model \(\mathcal {I}\), the MainOrchestrator initiates Stage 2. In this stage, the orchestrator performs knowledge-guided Feasibility Analysis to select a feasible exploit strategy \(\mathcal {S}\) from ranked candidates rather than directly instantiating a fixed payload template. The input to this decision process is the structured state produced in Stage 1, including the inferred VP capability (e.g., PC control, arbitrary read, arbitrary write), enabled mitigations or sandbox constraints (e.g., NX, Canary, PIE, RELRO, and Seccomp-like restrictions when inferable), architecture and calling convention, available program resources (e.g., PLT/GOT entries, writable sections, candidate gadgets, libc availability), interaction opportunities, local buffer bounds, and forbidden-byte constraints. The orchestrator first specifies the exploitation objective and estimates the Exploit Payload Requirements, namely the minimum memory footprint and prerequisite facts required by a candidate payload. Applying \(L_2\) reasoning rules (e.g., the Resource Constraint Analysis Framework), it compares these requirements against the concrete bounds extracted from \(\mathcal {I}\), such as controllable memory regions, writable storage, available output channels, and the number of interaction rounds. The resulting planning procedure has three operational steps. First, the orchestrator constructs a candidate set of exploit trajectories by activating the rule chains whose preconditions match the current state. For example, PC control with NX disabled activates a ret2shellcode candidate; PC control with NX enabled and callable libc resources activates ret2libc or ROP candidates; PIE, ASLR, or Canary constraints prepend leakage sub-goals; insufficient contiguous overwrite space activates stack-pivot or staged-write candidates; and a format-string arbitrary-write primitive activates GOT-overwrite candidates only when the relocation policy allows it. Second, each candidate is passed through hard feasibility gates derived from the current program facts. Examples include rejecting direct shellcode execution under NX, rejecting GOT overwrite under Full RELRO, rejecting direct return-address overwrite when Canary is enabled but no leak primitive exists, and rejecting absolute-address strategies when PIE/ASLR is active and no address disclosure channel is available. Third, the remaining candidates are ranked using an explicit feasibility rubric that favors satisfied preconditions, fewer unresolved measurements, lower implementation complexity, stronger interaction reliability, and consistency with the accumulated failure history. Thus, the \(L_2\) rule chains constrain and audit the search space, but they do not prescribe a one-to-one mapping from vulnerability type to exploit script. After selecting the highest-ranked viable trajectory, the orchestrator performs a Meta-Cognitive Audit to check whether the proposed path has a complete dependency chain from VP to EP and whether it violates common architectural pitfalls, such as omitting Canary preservation, using unmeasured offsets, ignoring stack alignment, or assuming a reusable interaction loop that \(\mathcal {I}\) does not support. The selected strategy is then decomposed into a Directed Acyclic Graph (DAG) of atomic sub-tasks. Each node is represented conceptually as \(\langle \textit{action}, \textit{preconditions}, required\_measurements , produced\_facts , \textit{consumers}\rangle\). For instance, a leak-first ret2libc plan may contain nodes for leaking a GOT entry, computing the libc base, measuring the overwrite offset, constructing the argument setup, and hijacking control; a stack-pivot plan may contain nodes for locating writable storage, measuring the pivot relation, preparing the secondary frame, and transferring the stack pointer. We denote this DAG as the strategic specification \(\mathcal {G_{S}}\), which determines the concrete measurement mandates in Stage 3 and the synthesis constraints in Stage 4. Stage 3: Measurementexpert-driven dynamic probing To resolve uncertainties caused by compiler-specific padding and complex memory alignments, the MeasurementExpert executes Stage 3 to dynamically measure Deterministic Memory Invariants. Guided strictly by the measurement mandates derived from \(\mathcal {G_{S}}\), the orchestrator dispatches a dynamic probing task (③) to the MeasurementExpert to complete concrete memory measurements via GDB. This validation process encompasses four analytical dimensions. The first is Deterministic Offset Calibration, which calculates precise relative distances (e.g., from an input buffer to the saved return address) using cyclic generation. The second is Stack Frame Measurement, which determines the relative layout between local variables for stack pivoting or localized overwrites. The third is Heap Layout Introspection, which dynamically analyzes chunk adjacency and binning conditions for metadata corruption. The fourth is Constraint-Specific Probing, which identifies vulnerability-unique invariants, such as format string argument indices. By synthesizing these runtime observations, the MeasurementExpert returns Grounded Measurements \(\mathcal {M}\) (e.g., exact relative offsets) back to the orchestrator as measurement results (④). This anchors the abstract strategy to the deterministic physical reality of the target environment before the system synthesizes the final exploit. Stage 4: Exploitcrafter-driven exploit code generation In the generation module, ExploitCrafter translates the abstract strategy \(\mathcal {S}\) into functional exploit code. At this stage, the orchestrator brings together the semantic interaction details (\(\mathcal {I}\)), the algorithmic sub-tasks (\(\mathcal {G_{S}}\)), and the grounded numerical measurements (\(\mathcal {M}\)) into a comprehensive parametric specification \(\Sigma\). This assembled specification then serves as the baseline context (\(\Sigma _0\)) for the subsequent validation stage. ExploitCrafter serves as a stateless code synthesizer. It receives \(\Sigma\) through a code-generation task dispatch (⑤) and translates the abstract constraints directly into a sequential pwntools API script. As an implementation agent, it focuses on three concrete requirements. It first ensures interaction synchronization by aligning the script’s I/O operations (e.g., recvuntil) with the target program using the exact prompt literals extracted in Stage 1. It then performs payload staging, organizing execution blocks for multi-stage attacks so that process continuity is preserved as specified by \(\Sigma\). It also enforces the required constraints, including architectural alignment requirements (e.g., 16-byte stack padding) and restrictions on invalid characters such as NULL bytes. The resulting artifact is a structured, executable Python exploit script (⑥). It includes absolute configuration constants and staged exploit logic, and is ready for testing in the verification module. Stage 5: MeasurementExpert-driven stratified feedback and remediation To address the challenge of coarse-grained execution feedback, the MainOrchestrator enters a closed-loop diagnostic stage after exploit code generation. Rather than relying solely on a binary pass-or-fail outcome, the system uses an iterative feedback loop in which MeasurementExpert systematically identifies the root causes of execution failures and guides subsequent remediation. During verification, the MainOrchestrator assigns evaluation of the generated script to MeasurementExpert. MeasurementExpert first runs the script against the target binary in a standard environment to assess the outcome. If the execution fails, for example by timing out or producing malformed output, MeasurementExpert captures the standard error logs and then moves into a debugging environment, such as GDB, to re-execute the exploit. This allows it to reconstruct the crash context faithfully by capturing register states and stack layouts at the exact point of failure. By consulting the diagnostic knowledge base, specifically the pitfall set \(\mathcal {P}\) corresponding to \(L_4\), MeasurementExpert maps observable telemetry errors to logical root causes. Table 3 lists representative failure-to-remediation mappings rather than an exhaustive diagnostic taxonomy. These findings are then returned as a DiagnosticReport that identifies whether the failure was caused by inaccurate parameters, implementation errors, or a fundamentally flawed strategy. The iterative remediation process is strictly stateful. To prevent the repeated execution of a failing approach, the orchestrator maintains a history of past failures (\(\textit{Hist}_{\textit{failures}}\)) within the execution context. By integrating the DiagnosticReport into this historical record, the orchestrator proceeds to one of three coarse-grained remediation categories. If the failure arises from numerical inaccuracies, the orchestrator returns to Stage 3 for parameter-level adjustment. It instructs MeasurementExpert to re-measure specific memory offsets through GDB and update the parameter state \(\mathcal {M}\). If the underlying logic remains sound but execution violates architectural constraints, such as unaligned stack pointers or forbidden characters, the orchestrator returns to Stage 4 for implementation-level refinement. It then instructs ExploitCrafter to regenerate the script under the revised structural constraints. If the core exploitation strategy is instead found to be invalid, the orchestrator initiates strategic backtracking and returns to the planning phase in Stage 2. It consults the failure history to prune the current path and proposes a new strategic trajectory, which is then subjected to dynamic re-measurement (Stage 3) before code generation is attempted again. This diagnostic loop continues until the orchestrator confirms successful end-to-end exploit execution based on a concrete success signal from the local execution environment, namely a stable interactive shell or successful flag retrieval. The loop also ends when a predefined iteration limit (\(N_{\textit{max}}\)) is reached. Rather than being hard-coded, \(N_{\textit{max}}\) is a configurable hyperparameter specified by human security experts. This allows the system to bound the exploration space according to target complexity and computational budget, preventing unbounded analysis loops and controlling the overall cost of LLM inference. By converting opaque execution crashes into structured, stateful feedback, this mechanism systematically repairs the generated exploit code and resolves the logical-physical discrepancies inherent in automated exploitation. The full end-to-end execution flow of this interactive generation process, together with its condition-driven backtracking behavior, is formally abstracted in Algorithm 1. As shown in Algorithm 1, the diagnostic loop begins by taking the initial exploit context \(\Sigma _0\) together with the explicit pitfall set \(\mathcal {P}\), and by initializing an empty failure sequence \(\textit{Hist}_{\textit{failures}}\) (Line 1). Subject to the attempt limit \(N_{\textit{max}}\), the system then enters a recurring generate-and-verify cycle (Line 2). In each iteration, ExploitCrafter turns the current context \(\Sigma\) into an executable exploit candidate \(\mathcal {E}\) (Line 3). Functional evaluation of this candidate against the target binary is then assigned to MeasurementExpert (Line 4). If the payload obtains a shell or reads the flag, the expert returns a success signal, and the iterative loop terminates successfully (Lines 5–6). If the initial execution fails, MeasurementExpert moves into a debugging environment to reconstruct the full crash context. By analyzing this context and matching telemetry symptoms against the heuristic pitfall set \(\mathcal {P}\), the expert produces a structured failure diagnostic report. This report is immediately appended to the global trajectory history \(\textit{Hist}_{\textit{failures}}\) (Line 7). Guided by the root-cause classification in the report, the orchestrator routes the loop into one of three coarse-grained remediation branches. A ParamError triggers a lightweight fallback to Stage 3, where dynamic re-probing is guided by the specific \(\textit{report}\) to update memory offsets without discarding the established strategic trajectory \(\Sigma\) (Lines 8–10). An ImplError triggers a structural patch in Stage 4, directly refining the constraints of the current exploit context \(\Sigma\) to fix superficial issues such as bad characters or stack misalignment (Lines 11–12). A fundamental StrategyError, by contrast, triggers a deeper cognitive rollback. The orchestrator then returns to Stage 2 for global replanning and synthesizes an alternative path based on both the original program context \(\Sigma\) and the accumulated failures in \(\textit{Hist}_{\textit{failures}}\) (Line 14). The system then requires dynamic re-probing of this new strategic trajectory in Stage 3 to resolve its specific memory constraints, ensuring that it is fully grounded before code generation begins (Lines 15–16). If the attempt budget is exhausted before shell access or flag retrieval is observed, the state machine terminates gracefully and returns a failure (Line 20). Implementation This work builds a prototype of PwnAgent based on the LangGraph framework, orchestrating the task flow and state evolution among multiple agents through a State Graph. The system is developed using Python, with the core logic modules and the hierarchical knowledge base cumulatively containing approximately 23,400 lines of code. Below, we describe the physical instantiation of the system in binary analysis, strategy orchestration, and dynamic verification. Program perception This module provides the underlying support that enables ProgramAnalyzer to perform static layout modeling. It integrates checksec, file, and binary parsing interfaces based on readelf. The system also extracts disassembly and pseudocode representations of the target program using IDA Pro deployed in a Docker container. To support exploit-resource discovery, we integrate ROPgadget to extract code-reuse gadgets from the binary and use one_gadget to identify exploitation constraints in C library functions. During analysis, this layer converts the raw output of low-level analysis tools into structured ProgramInfo objects for logical planning by the upper-level orchestrator. Cooperative orchestrator This component implements the scheduling logic of MainOrchestrator. Its core design builds on LangGraph’s AgentState mechanism to maintain global state and coordinate the three specialized agents. We define and enforce a structured interaction protocol based on JSON Schema, which allows agents to pass intermediate data objects directly, including ProgramInfo, Strategy, Measurements, and Diagnostics, thereby reducing the information loss and semantic ambiguity that may arise from natural-language paraphrasing. To improve reliability when analyzing long execution paths, the system uses a persistent storage mechanism (Checkpointing) that supports logical backtracking across analysis nodes. Dynamic measurement and verifier This module is driven primarily by MeasurementExpert. It uses GDB for dynamic state extraction and integrates the pwntools framework to manage interaction with the execution environment. To address the uncertainty of static analysis in determining stack overflow offsets, we implement pattern-based probing primitives using cyclic sequences together with exception register states captured by GDB, which enables precise identification of stack overflow offsets. The final exploit code is executed in an Ubuntu 18.04 LTS environment, and the system verifies its effectiveness by confirming successful execution of a validation command through the obtained shell or successful retrieval of the flag. For reproducibility, Appendix A presents a simplified design of the structured system prompt templates that guide these agents. Evaluation In this section, we introduce the evaluation benchmark established for the AEG framework and discuss the evaluation results. We address the following research questions in our evaluation: RQ1 End-to-end effectiveness. What is the overall success rate of automatically completing the entire exploit generation process, and how does this rate vary across backend LLMs, difficulty levels, vulnerability classes, and CPU architectures? How do PwnAgent and PwnGPT compare on the public PwnGPT benchmark? RQ2 Completion degree. To what extent can individual exploit sub-tasks be completed automatically? RQ3 Design choice evaluation. What is the impact of knowledge-base-guided reasoning, static-dynamic synergistic analysis, and feedback-driven optimization on the performance of PwnAgent? Evaluation setup Benchmark Dataset We constructed a curated binary exploitation benchmark to measure the performance of our AEG framework under reproducible CTF-style constraints. The dataset comprises 66 target programs from established capture-the-flag platforms, including BUUCTF, XCTF, and CTFshow. Appendix B reports the benchmark metadata used for characterization, with a dedicated field for source-platform provenance together with challenge types, enabled mitigations, and intended exploit methods. These platforms supply the supporting runtime environments and public validation materials used by the evaluators to replay each task in a controlled local environment. For each benchmark task, the agent-facing input follows the input assumption in Section Problem Scoping and Threat Model: the system receives only the target binary and the challenge-provided libc when such a library is distributed. The platform name, vulnerability class, difficulty label, security-measure summary, and intended exploit method reported in Appendix B are used only for dataset characterization and post-hoc evaluation. They are not included in the prompts, not exposed through the workspace, and not provided as auxiliary hints. Likewise, no challenge description, source code, debug symbols, vulnerability location, proof-of-concept input, crash-triggering input, reference exploit, or write-up is provided to PwnAgent during evaluation. The benchmark tasks and their public solution materials are also excluded from the external KB and from all role-specific prompt shards. During evaluation, PwnAgent is not provided with any online write-up retrieval interface or benchmark-specific solution source. We distinguish generic technique overlap from task leakage: the KB may contain general knowledge about common exploitation techniques, but it excludes materials that describe the specific benchmark instances. The benchmark design follows three practical criteria. First, each target must be a Linux ELF executable that can run reliably in our Ubuntu 18.04 LTS testing environment. Second, the task should require at least one concrete exploitation capability, such as offset calibration, information leakage, address recovery, ROP-chain construction, stack pivoting, or interaction synchronization. Third, the overall set should provide a gradual difficulty progression rather than only trivial or extremely specialized instances. These criteria lead to an intentionally imbalanced but capability-oriented dataset. Most targets are x86 binaries and most vulnerability instances are stack overflows, because this setting supports controlled variation over mitigations (e.g., NX, Canary, PIE, RELRO) and exploit techniques while keeping the execution environment reproducible. We also include format-string, heap-corruption, integer-overflow, ARM, and MIPS cases as limited probes beyond the dominant stack/x86 setting. However, these minority categories are not large enough to support broad claims about balanced vulnerability-class or architecture coverage. Because open platforms typically do not provide uniform difficulty metrics, we manually assigned difficulty labels using an operational rubric applied to a validated exploit path for each challenge. The rubric considers three observable factors: (1) exploit-chain depth, including interdependent stages such as information leakage, address recovery, stack pivoting, ROP-chain construction, or heap shaping; (2) constraint burden, including mitigations and architectural constraints that must be handled by the exploit; and (3) interaction and state complexity, ranging from single-round input to multi-stage synchronization or stateful menu-driven manipulation. Using this rubric, Easy challenges admit a short exploit chain with simple interaction; Medium challenges require at least one critical intermediate stage or substantive constraint-handling step; and Hard challenges require multiple interdependent stages, complex state maintenance, or the joint handling of multiple constraints. Two experts with practical experience in vulnerability exploitation independently labeled all 66 challenges using this rubric. Before adjudication, they assigned the same label to 60/66 challenges (90.91%). For challenges with disagreement, a third expert independently assigned the final labels using the same rubric. As summarized in Table 4, the evaluation dataset includes 66 configurations categorized by vulnerability class and difficulty level, with 12 easy, 40 medium, and 14 hard challenges. Rather than claiming a representative sample of the full pwn ecosystem, this center-heavy distribution is intended to create a discriminative testbed for the specific capabilities studied in this paper: runtime measurement, strategy refinement, and feedback-driven exploit repair. Experimental setup and model configuration Our evaluation experiments were conducted on an Ubuntu 18.04 LTS virtual machine provisioned with 2 CPU cores and 16 GB of RAM. The core system was implemented in Python and integrated GDB for auxiliary dynamic analysis. To isolate the environment, we deployed the static analysis backend, IDA Pro, in a Docker container running inside the virtual machine. For the LLM component, our main evaluation uses DeepSeek-V3.1-Terminus and GLM−4.6 as the foundation models for the AEG framework. To examine whether the system-level gain persists under a recent coding- and agent-oriented backend, we also include Kimi-K2.6. For each backend model, PwnAgent and PwnGPT are evaluated on the same benchmark under the same agent-facing inputs and success criterion. For compact table headers, DeepSeek denotes DeepSeek-V3.1-Terminus. Metrics Following PwnGPT’s staged decomposition of AEG capabilities, but adopting stricter execution-based criteria for final exploit success, we deconstructed the exploit generation pipeline into four sequential stage capabilities. Critical Information Analysis is credited when the system correctly extracts the target architecture, binary attributes, enabled security mechanisms, and relevant interaction constraints. VP Identification is credited when it identifies at least one real vulnerability primitive and its exploitable code location, rather than merely naming a generic vulnerability class. Exploit Strategy Planning is credited only when the generated strategy contains at least one complete and feasible exploit chain from the identified VP to the final exploitation objective. A task may admit multiple valid exploitation paths; therefore, the strategy is counted as correct if one path is complete, technically sound, and consistent with the observed mitigations and resources, including required leakage, address recovery, mitigation bypass, gadget/resource preparation, and staged interaction steps. High-level ideas, incomplete chains, or strategies that violate constraints such as NX, Canary, PIE/ASLR, RELRO, or calling-convention requirements are not counted as successful planning. Finally, EP Materialization is credited only when the system generates an executable exploit script that runs in the local evaluation environment and stably obtains an interactive shell or reads the flag. Scripts with minor syntax errors, missing constants, incorrect offsets, incomplete payloads, or only a plausible payload skeleton are counted as failures at this stage. The first three criteria are assessed by expert audit of structured outputs and available execution traces, whereas EP Materialization is determined by direct replay of the generated script; expert review at this stage is used only to attribute the failure mode, not to award partial credit. For the end-to-end metric, a task is counted as successful under the same strict criterion as EP Materialization: the system must automatically produce a runnable exploit script that obtains shell access or retrieves the flag in the local environment. Intermediate achievements such as identifying a VP, controlling the program counter, leaking an address, or constructing an EP are recorded only in the stage-level metrics and are not counted as end-to-end successes unless they lead to shell access or flag retrieval. Because these testing metrics form a strict dependency chain, an error at any intermediate stage prevents the task from being counted as an end-to-end success. Effectiveness of the entire framework We evaluate PwnAgent using its success rate on the end-to-end exploit generation task. To separate the contribution of the framework from the underlying capabilities of the language models, we compare PwnAgent with PwnGPT on the full benchmark under identical backend-model settings. As summarized in Table 5, PwnAgent achieves an overall success rate of 33.33% under DeepSeek-V3.1-Terminus, compared with 19.70% for PwnGPT, corresponding to an absolute improvement of 13.64 percentage points. Under GLM−4.6, PwnAgent achieves 34.85%, compared with 18.18% for PwnGPT, yielding an improvement of 16.67 percentage points. To examine whether the system-level gain persists under a more recent coding-oriented backend, we further evaluate both systems using Kimi-K2.6. Under this backend, PwnGPT reaches 31.82%, approaching the 33.33% and 34.85% achieved by PwnAgent under DeepSeek-V3.1-Terminus and GLM−4.6, respectively. When paired with the same Kimi-K2.6 backend, PwnAgent improves over PwnGPT from 31.82% to 62.12%, corresponding to a 30.30 percentage-point gain. These paired comparisons suggest that the contribution of PwnAgent remains material as the capability of the underlying model improves. Figure 3 further decomposes the end-to-end results by difficulty. Under DeepSeek-V3.1-Terminus, PwnAgent improves the success rate from 66.67% (8/12) to 91.67% (11/12) on Easy tasks and from 12.50% (5/40) to 25.00% (10/40) on Medium tasks. Under GLM−4.6, the corresponding improvements are from 66.67% (8/12) to 91.67% (11/12) and from 10.00% (4/40) to 27.50% (11/40). Under Kimi-K2.6, the Easy subset is already close to saturation for both systems, with success rates of 91.67% (11/12) for PwnGPT and 100.00% (12/12) for PwnAgent. The largest gain instead occurs in the Medium category, where PwnAgent improves from 22.50% (9/40) to 60.00% (24/40). The Hard-category success rate also increases from 7.14% (1/14) to 35.71% (5/14). Thus, the improvement under Kimi-K2.6 is not confined to Easy tasks. Nevertheless, the lower success rate on Hard tasks indicates that heavily constrained and multi-stage exploitation remains challenging. Table 6 reports the end-to-end results by vulnerability class and CPU architecture for the evaluated backends. For stack-overflow tasks, PwnAgent solves 16/46 challenges under DeepSeek-V3.1-Terminus and 18/46 under GLM−4.6, compared with 9/46 for PwnGPT under both backends. For format-string tasks, both systems solve 2/9 challenges under both backends. PwnAgent solves 1/7 heap-corruption challenges under both backends, while PwnGPT solves none. For integer-overflow tasks, PwnAgent solves 3/4 and 2/4 challenges under DeepSeek-V3.1-Terminus and GLM−4.6, respectively, compared with 2/4 and 1/4 for PwnGPT. Under Kimi-K2.6, PwnAgent solves 29/46 stack-overflow tasks, 5/9 format-string tasks, 3/7 heap-corruption tasks, and 4/4 integer-overflow tasks. The corresponding results for PwnGPT are 14/46, 3/9, 2/7, and 2/4, respectively. These results indicate that the observed gain is concentrated primarily in the stack-overflow subset. Because the remaining vulnerability classes contain substantially fewer tasks, their results should be interpreted as limited probes rather than stable estimates of vulnerability-class-specific performance. The architecture breakdown leads to a similar conclusion. The end-to-end totals include the ARM and MIPS challenges. On the dominant x86/x86-64 subset, PwnAgent solves 21/63 challenges under DeepSeek-V3.1-Terminus and 22/63 under GLM−4.6, compared with 13/63 and 12/63 for PwnGPT. PwnAgent solves one of the two ARM challenges under both backends, whereas neither system solves the single MIPS challenge. Under Kimi-K2.6, PwnAgent solves 40/63 x86/x86-64 tasks, 1/2 ARM tasks, and 0/1 MIPS tasks. The corresponding results for PwnGPT are 21/63, 0/2, and 0/1, respectively. Given the small size of the non-x86 subsets, these results should be treated as limited cross-architecture probes rather than evidence of broad architecture coverage. Taken together, the results suggest that the structured pipeline, hierarchical knowledge guidance, and dynamic feedback loop help the underlying models convert static program observations into executable exploits. However, the low absolute success rate on Hard tasks continues to highlight the intrinsic difficulty of end-to-end automated exploit generation. Analysis of the failed cases suggests two primary causes, both of which reflect the current limitations of language models. The first is Long-Horizon Strategy Hallucination (Strategy Failure). In complex challenges that require multi-stage vulnerability chaining, such as combining an information leak with heap manipulation to bypass PIE and Canary, the models struggle to sustain extended deductive reasoning. Even with knowledge-base guidance, the orchestrator sometimes selects an exploitation path that appears theoretically plausible but does not fit the actual environment, leading to a fundamentally flawed strategy that cannot be executed. The second is Precision Loss in EP Materialization (Implementation Failure). Even when the system identifies the correct high-level strategy, translating that semantic plan into an executable Python script still requires byte-level precision. Failures at this stage usually stem from intricate physical constraints. Typical examples include subtle I/O buffering mismatches in interactive prompts, complex payload formatting requirements (e.g., bad character evasion, p64 versus p32 packing), and uninformative crashes inside C library functions. When execution failures produce ambiguous diagnostic signals, such as unpredictable timeouts or missing crash contexts, the feedback-driven self-correction mechanism has difficulty identifying the root cause and may remain stuck in a retry loop until the iteration limit is reached. Evaluation on the PwnGPT public benchmark For direct comparison with prior work, we also evaluate PwnAgent and PwnGPT on the public 19-task benchmark released with PwnGPT (Peng et al. 2025). Both systems retain their respective system prompts and workflow configurations from the main evaluation, and the KB used by PwnAgent remains unchanged. The agent-facing input policy, execution environment, and strict shell-or-flag success criterion are also unchanged, and both systems use the same backend model in each comparison. No benchmark metadata or public solution material is provided to either system. Table 7 reports the paired end-to-end results under DeepSeek-V3.1-Terminus, GLM−4.6, and Kimi-K2.6. Under DeepSeek-V3.1-Terminus, PwnAgent solves 10/19 tasks (52.63%), whereas PwnGPT solves 6/19 (31.58%), a difference of 21.05 percentage points. Under GLM−4.6, PwnAgent solves 9/19 tasks (47.37%), whereas PwnGPT solves 4/19 (21.05%), a difference of 26.32 percentage points. Under Kimi-K2.6, PwnAgent solves 18/19 tasks (94.74%), whereas PwnGPT solves 13/19 (68.42%), also a difference of 26.32 percentage points. Under all three matched backend models, PwnAgent outperforms PwnGPT on the public benchmark, providing cross-benchmark evidence from a distinct task collection. Completion degree of AEG To analyze the sources of PwnAgent’s performance gains over the baseline, we conducted a detailed comparative analysis of PwnAgent and PwnGPT using the four stage-level indicators defined in Section Evaluation Setup. For this part of the evaluation, we selected DeepSeek-V3.1-Terminus as the representative base model. To ensure an objective and fair assessment, we invited two senior offensive and defensive security experts to independently audit and evaluate the intermediate outputs of the systems. A third expert served as an arbitrator to resolve any disagreements. The experimental results shown in Fig. 4 report the task completion status of both systems at each critical node. In the early stages, such as critical information analysis and VP identification, PwnAgent maintains stable performance. By contrast, PwnGPT already shows early performance degradation during VP identification, with a success rate of 95.45%. In the subsequent Exploit Strategy Planning stage, which maps VPs to EPs under multiple protection mechanisms, the success rate of PwnAgent remains at 50.00%, while that of PwnGPT drops to 33.33%. In the final EP Materialization stage, PwnAgent achieves an end-to-end exploit generation rate of 33.33%, higher than the 19.70% baseline achieved by PwnGPT. These stage-by-stage results suggest that the performance gain of PwnAgent does not come from an isolated breakthrough at a single stage, but from improved execution robustness across the critical nodes in the pipeline. Ablation study To evaluate the contribution of each component, we developed three ablation variants of the full architecture. In the No-KB variant, we disable the external exploit knowledge base. Each agent relies solely on the internal knowledge of the underlying language model and a limited set of general prompts. The remainder of the architecture, including multi-agent partitioning and the invocation of analysis tools, remains unchanged. In the No-DA variant, we disable runtime parameter measurement and feedback-guided optimization based on Diagnostics. The system relies exclusively on the static ProgramInfo generated by the ProgramAnalyzer and executes the Strategy and Exploit in a single-round attempt without constructing Measurements or performing iterative corrections. In the No-KB+No-DA variant, both external knowledge guidance and dynamic analysis are disabled, leaving only static program analysis, role-specialized orchestration, and single-round exploit generation. We evaluated the full system and all variants on the identical benchmark set. The end-to-end success rates are shown in Table 8. Table 8 shows that, with DeepSeek-V3.1-Terminus as the backend model, the full system achieves a success rate of 33.33%. By comparison, the No-KB variant drops to 10.61%, and the No-DA variant drops to 12.12%. With GLM−4.6, the full system reaches 34.85%, while the No-KB and No-DA variants fall to 7.58% and 15.15%, respectively. Under Kimi-K2.6, the full system achieves 62.12%, whereas the No-KB and No-DA variants achieve 33.33% and 42.42%, respectively. The No-KB+No-DA variant provides a stricter ablation setting for measuring the residual capability of static analysis, role-specialized orchestration, and single-round exploit generation after both proposed components are removed. Under this configuration, the success rate decreases further to 7.58% with DeepSeek-V3.1-Terminus, 6.06% with GLM−4.6, and 24.24% with Kimi-K2.6. For each backend, the No-KB+No-DA success rate is lower than the rates of both single-component ablations. This pattern indicates that KB guidance and dynamic analysis each make an incremental contribution beyond the retained static analysis and role-specialized orchestration pipeline. The larger drop in No-KB than in No-DA further suggests that, for this task distribution, the knowledge base contributes more to end-to-end performance. A closer analysis by protection configuration and structural complexity provides further insight. On instances with relatively simple structures and weak protections, the performance gap among the configurations is small. On instances that require information leakage and multi-stage ROP chain construction under multiple active protections, however, the success rate of the No-DA variant decreases markedly. In these more complex settings, common failure patterns include stack overflow offset deviations, format string index errors caused by changes in variable arrangement, and illegal instruction exceptions triggered by unaligned ROP chains. These failures directly reflect the Introspection Deficit (C1) and Feedback Ambiguity (C3) challenges. Without obtaining precise Measurements through MeasurementExpert or using Diagnostics for targeted correction, the system cannot bridge the Logical-Physical Gap. The failures of the No-KB variant arise primarily at the strategy and logic levels, reflecting the Knowledge Representation Gap (C2). Even when static analysis provides relatively complete ProgramInfo, this configuration lacks the structured domain knowledge required for offensive security. As a result, the agents often select exploit techniques that are incompatible with the active memory protections or fail to account for necessary intermediate stages, such as obtaining an information leak before address recovery. These omissions lead to exploit chains that are not viable in practice. With the hierarchically organized KB and its associated CoT structure, MainOrchestrator can systematically evaluate the available capabilities and constraints at the strategic level. It explicitly represents intermediate sub-goals, such as information leakage, address recovery, and resource preparation, within the Strategy, which improves the success rate in complex scenarios. The results confirm the distinct roles of the different components. The knowledge base addresses C2 by improving the reliability of high-level exploit strategy planning and multi-stage chain construction. At the same time, the static-dynamic synergistic analysis and feedback-guided optimization directly address C1 and C3 by providing precise runtime parameters and enabling iterative correction of implementation-level errors. Together, these components constitute the main sources of the performance gains achieved by the proposed framework. Limitations Benchmark representativeness This evaluation is based on 66 target programs selected from mainstream competition platforms such as BUUCTF, XCTF, and CTFshow. The dataset is intentionally capability-oriented and is not a balanced sample of the full pwn ecosystem. In particular, 46 of the 66 tasks are stack-overflow challenges, while only 7 are heap-corruption challenges, 9 are format-string challenges, and 4 are integer-overflow challenges. The architecture distribution is also dominated by x86 binaries; the benchmark contains only two ARM cases and one MIPS case. These ARM/MIPS tasks are useful as limited probes of cross-architecture behavior, but they are insufficient for drawing broad conclusions about non-x86 exploitation. Although the benchmark covers diverse mitigation mechanisms and combinations of typical vulnerability patterns, there are still significant differences between CTF challenges and real-world industrial-grade software in terms of code scale, dependency complexity, and input vector diversity. Therefore, the empirical results should be interpreted as evidence for PwnAgent’s behavior on this curated, x86-dominated CTF benchmark rather than as a general claim about all binary exploitation scenarios. Data contamination We distinguish benchmark-specific system leakage from model-level parametric contamination. For the former, we exclude benchmark-specific artifacts from the external KB, role-specific prompt shards, and agent-facing inputs, including public write-ups, reference exploits, proof-of-concept inputs, solution scripts, flags, and challenge-specific analysis notes. This procedure reduces the risk that PwnAgent solves a task by consuming task-specific external materials during evaluation. However, it cannot fully eliminate parametric contamination in the underlying LLMs, a recognized limitation of LLM benchmark evaluation where public evaluation materials may overlap with pre-training or instruction-tuning corpora (Xu et al. 2024; Li et al. 2024). In our setting, this means public CTF challenges and their write-ups may have appeared in large-scale model-training data. Moreover, both the KB source corpus and the benchmark draw heavily from the public CTF ecosystem. Shared exploitation patterns, techniques, and conventions may therefore favor in-domain tasks, and excluding task-specific materials does not remove this broader ecosystem overlap. Therefore, the results should be interpreted as performance under a task-disjoint external KB and controlled agent-facing inputs, rather than as evidence of complete model-level decontamination or generalization beyond the public CTF ecosystem. CoT faithfulness PwnAgent uses CoT-style inquiries, rule chains, and reference trajectories as an engineering mechanism for structuring domain knowledge and constraining exploit planning. However, prior work has shown that generated CoT explanations are not necessarily faithful representations of a model’s internal decision process (Turpin et al. 2023; Lanham et al. 2023). We therefore do not claim that the CoT traces produced or consumed by PwnAgent reveal the true latent reasoning of the underlying LLM. Our evaluation measures externally observable behavior, including exploit success, runtime measurements, diagnostic corrections, and final shell/flag retrieval, rather than CoT faithfulness itself. Future work could evaluate faithfulness directly or combine the KB with more formal execution or symbolic checks for individual reasoning steps. Conclusion We present PwnAgent, an LLM-driven multi-agent framework for end-to-end AEG. By addressing three core challenges, namely limited runtime introspection, insufficient domain knowledge, and ambiguous feedback, PwnAgent establishes a structured Sense-Model-Diagnose synergy. This architecture bridges the Logical-Physical Gap by turning abstract VPs into functional EPs grounded in physical memory. Evaluation on a curated, difficulty-graded 66-task pwn benchmark shows that PwnAgent outperforms the evaluated PwnGPT baseline under the same model settings, benefiting from structured knowledge guidance and an execution-grounded measurement-and-repair loop. At the same time, the x86- and stack-heavy benchmark distribution and the remaining low absolute success rate indicate that autonomous binary exploitation remains an open problem, particularly for heap-oriented, non-x86, and long-horizon exploitation scenarios. Data availability The data that support the findings of this study are available from the corresponding author upon reasonable request. Material availability Not applicable. Code availability Project name: PwnAgent Programming language: Python (based on LangGraph framework) Operating system(s): Linux License: Not applicable (available upon reasonable request). References Avgerinos T, Cha SK, Hao BLT, Brumley D (2011) AEG: Automatic exploit generation. In: Proceedings of the Network and Distributed System Security Symposium (NDSS 2011) Babaey V, Ravindran A (2025) Genxss: An ai-driven framework for automated detection of xss attacks in wafs. SoutheastCon 2025. IEEE, pp 1519–1524 Bao T, Wang R, Shoshitaishvili Y, Brumley D (2017) Your exploit is mine: Automatic shellcode transplant for remote exploits. In: 2017 IEEE Symposium on Security and Privacy (SP), pp. 824–839 https://doi.org/10.1109/SP.2017.67 Brumley D, Poosankam P, Song D, Zheng J (2008) Automatic patch-based exploit generation is possible: Techniques and implications. In: 2008 IEEE Symposium on Security and Privacy (sp 2008), pp. 143–157 https://doi.org/10.1109/SP.2008.17 Cha SK, Avgerinos T, Rebert A, Brumley D (2012) Unleashing mayhem on binary code. In: 2012 IEEE Symposium on Security and Privacy, pp. 380–394 https://doi.org/10.1109/SP.2012.31 Chen W, Zou X, Li G, Qian Z (2020)Koobe: towards facilitating exploit generation of kernel out-of-bounds write vulnerabilities. Proceedings of the 29th USENIX Conference on Security Symposium. SEC’20. USENIX Association. USA, Chen J, Lin H, Han X, Sun L (2024) Benchmarking large language models in retrieval-augmented generation. AAAI’24/IAAI’24/EAAI’24 https://doi.org/10.1609/aaai.v38i16.29728 Chen X, Lin M, Schärli N, Zhou D (2024) Teaching large language models to self-debug. International Conference on Learning Representations.(ICLR 2024) Deng G, Liu Y, Mayoral-Vilches V, Liu P, Li Y, Xu Y, Zhang T, Liu Y, Pinzger M, Rass S (2024) Pentestgpt: evaluating and harnessing large language models for automated penetration testing. In: Proceedings of the 33rd USENIX Conference on Security Symposium. SEC ’24. USENIX Association, USA Fang R, Bindu R, Gupta A, Zhan Q, Kang D (2024) LLM agents can autonomously hack websites. arXiv preprint arXiv:2402.06664https://doi.org/10.48550/arXiv.2402.06664 Garmany B, Stoffel M, Gawlik R, Koppe P, Blazytko T, Holz T (2018) Towards automated generation of exploitation primitives for web browsers. In: Proceedings of the 34th Annual Computer Security Applications Conference. ACSAC ’18, pp. 300–312. Association for Computing Machinery, New York, NY, USA https://doi.org/10.1145/3274694.3274723 Heelan S, Melham T, Kroening D (2019) Gollum: Modular and greybox exploit generation for heap overflows in interpreters. In: Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security. CCS ’19, pp. 1689–1706. Association for Computing Machinery, New York, NY, USA https://doi.org/10.1145/3319535.3354224 Hu H, Chua ZL, Adrian S, Saxena P, Liang Z (2015) Automatic generation of data-oriented exploits. In: Proceedings of the 24th USENIX Conference on Security Symposium. SEC’15, pp. 177–192. USENIX Association, USA Huang S-K, Huang M-H, Huang P-Y, Lai C-W, Lu H-L, Leong W-M (2012) Crax: Software crash analysis for automatic exploit generation by modeling attacks as symbolic continuations. In: Proceedings of the 2012 IEEE Sixth International Conference on Software Security and Reliability. SERE ’12, pp. 78–87. IEEE Computer Society, USA https://doi.org/10.1109/SERE.2012.20 Ispoglou KK, AlBassam B, Jaeger T, Payer M (2018) Block oriented programming: Automating data-only attacks. In: Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security. CCS ’18, pp. 1868–1882. Association for Computing Machinery, New York, NY, USA https://doi.org/10.1145/3243734.3243739 Jin B, Yoon J, Han J, Arik SO (2024) Long-context LLMs meet RAG: overcoming challenges for long inputs in RAG Kaniewski S, Schmidt F, Enzweiler M, Menth M, Heer T (2025) A systematic literature review on detecting software vulnerabilities with large language models. arXiv preprint arXiv:2507.22659https://doi.org/10.48550/arXiv.2507.22659 Khan S (2024) Ll-xss: End-to-end generative model-based xss payload creation. In: 2024 21st Learning and Technology Conference (L&T), pp. 121–126 https://doi.org/10.1109/LT60077.2024.10469151 Klischies D, Mackensen P, Moonsamy V (2025) Vulnerability where art thou? an investigation of vulnerability management in android smartphone chipsets. NDSS Symposium. https://doi.org/10.14722/ndss.2025.241161 Kurmus A, Mambretti A, Sorniotti A, Lenders V, Pfammatter D, Tellenbach B (2025) SoK: Automating kernel vulnerability discovery and exploit generation. 19th USENIX WOOT Conference on Offensive Technologies (WOOT ’25). pp 283–302 Lanham T, Chen A, Radhakrishnan A, Steiner B, Denison C, Hernandez D, Li D, Durmus E, Hubinger E, Kernion J (2023) Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702 Leverett,É (2025) Vulnerability Forecast: Year-End Review. In: Forum of Incident Response and Security Teams (FIRST) (2025). Accessed 9 March 2026. https://www.first.org/blog/20251229-Vulnerability-Forecast-Review Li Y, Guo Y, Guerin F, Lin C (2024) An open-source data contamination report for large language models. In: Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 528–541 Liu J, An H, Li J, Liang H (2022) Detecting exploit primitives automatically for heap vulnerabilities on binary programs. https://doi.org/10.48550/arXiv.2212.13990 Liu NF, Lin K, Hewitt J, Paranjape A, Bevilacqua M, Petroni F, Liang P (2024) Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics 12:157–173. https://doi.org/10.1162/tacl_a_00638 Meng R, Mirchev M, Böhme M, Roychoudhury A (2024) Large language model guided protocol fuzzing. Network and Distributed System Security (NDSS) Symposium. p 2024. https://doi.org/10.14722/ndss.2024.24556 Peng W, Ye L, Du X, Zhang H, Zhan D, Zhang Y, Guo Y, Zhang C (2025) PwnGPT: Automatic exploit generation based on large language models. In: Che, W., Nabende, J., Shutova, E., Pilehvar, M.T. (eds.) Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11481–11494. Association for Computational Linguistics, Vienna, Austria https://doi.org/10.18653/v1/2025.acl-long.562 Qian C, Liu W, Liu H, Chen N, Dang Y, Li J, Yang C, Chen W, Su Y, Cong X, Xu J, Li D, Liu Z, Sun M (2024) ChatDev: Communicative agents for software development. In: Ku, L.-W., Martins, A., Srikumar, V. (eds.) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15174–15186. Association for Computational Linguistics, Bangkok, Thailand https://doi.org/10.18653/v1/2024.acl-long.810 Ray S, Pan R, Gu Z, Du K, Feng S, Ananthanarayanan G, Netravali R, Jiang J (2025) Metis: fast quality-aware rag systems with configuration adaptation. Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles. pp 606–622 Shao M, Jancheska S, Udeshi M, Dolan-Gavitt B, Xi H, Milner K, Chen B, Yin M, Garg S, Krishnamurthy P, Khorrami F, Karri R, Shafique M (2024) Nyu ctf bench: a scalable open-source benchmark dataset for evaluating llms in offensive security. Proceedings of the 38th International Conference on Neural Information Processing Systems. NIPS ’24. Curran Associates Inc, Red Hook, NY, USA, Shao M, Xi H, Rani N, Udeshi M, Putrevu VSC, Milner K, Dolan-Gavitt B, Shukla SK, Krishnamurthy P, Khorrami F, Karri R, Shafique M (2025) CRAKEN: Cybersecurity LLM agent with knowledge-based execution. arXiv preprint arXiv:2505.17107https://doi.org/10.48550/arXiv.2505.17107 Shinn N, Cassano F, Gopinath A, Narasimhan K, Yao S (2023) Reflexion: language agents with verbal reinforcement learning. Proceedings of the 37th International Conference on Neural Information Processing Systems. NIPS ’23. Curran Associates Inc, Red Hook, NY, USA, Turpin M, Michael J, Perez E, Bowman S (2023) Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. Adv Neural Inf Process Syst 36:74952–74965 Wang L, Ma C, Feng X, Zhang Z, Yang H, Zhang J, Chen Z, Tang J, Chen X, Lin Y et al (2024) A survey on large language model based autonomous agents. Front Comp Sci 18(6):186345. https://doi.org/10.1007/s11704-024-40231-1 Wang Y, Zhang C, Xiang X, Zhao Z, Li W, Gong X, Liu B, Chen K, Zou W (2018) Revery: From proof-of-concept to exploitable. In: Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security. CCS ’18, pp. 1914–1927. Association for Computing Machinery, New York, NY, USA https://doi.org/10.1145/3243734.3243847 Wang D, Zhou G, Chen L, Li D, Miao Y (2024) Prophetfuzz: Fully automated prediction and fuzzing of high-risk option combinations with only documentation via large language model. In: Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. CCS ’24, pp. 735–749. Association for Computing Machinery, New York, NY, USA https://doi.org/10.1145/3658644.3690231 Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, Chi EH, Le QV, Zhou D (2022) Chain-of-thought prompting elicits reasoning in large language models. Proceedings of the 36th International Conference on Neural Information Processing Systems. NIPS ’22. Curran Associates Inc, Red Hook, NY, USA, Wu W, Chen Y, Xu J, Xing X, Gong X, Zou W (2018) Fuze: towards facilitating exploit generation for kernel use-after-free vulnerabilities. In: Proceedings of the 27th USENIX Conference on Security Symposium. SEC’18, pp. 781–797. USENIX Association, USA Wu Q, Bansal G, Zhang J, Wu Y, Li B, Zhu E, Jiang L, Zhang X, Zhang S, Liu J et al.: (2024) Autogen: Enabling next-gen llm applications via multi-agent conversations. In: First Conference on Language Modeling Xia CS, Paltenghi M, Le Tian J, Pradel M, Zhang L (2024) Fuzz4all: Universal fuzzing with large language models. In: Proceedings of the IEEE/ACM 46th International Conference on Software Engineering. ICSE ’24. Association for Computing Machinery, New York, NY, USA https://doi.org/10.1145/3597503.3639121 Xiao Z, Wang Q, Li Y, Chen S (2025) Prompt to pwn: Automated exploit generation for smart contracts. arXiv preprint arXiv:2508.01371https://doi.org/10.48550/arXiv.2508.01371 Xu J, Stokes JW, McDonald G, Bai X, Marshall D, Wang S, Swaminathan A, Li Z (2024) Autoattacker: A large language model guided system to implement automatic cyber-attacks. arXiv preprint arXiv:2403.01038https://doi.org/10.48550/arXiv.2403.01038 Xu C, Guan S, Greene D, Kechadi M (2024) Benchmark data contamination of large language models: A survey. arXiv preprint arXiv:2406.04244 Yang G, Zhou Y, Chen X, Zhang X, Han T, Chen T (2023) Exploitgen: Template-augmented exploit code generation based on codebert. J Syst Softw 197:111577. https://doi.org/10.1016/j.jss.2022.111577 Yang J, Prabhakar A, Narasimhan K, Yao S (2023) Intercode: standardizing and benchmarking interactive coding with execution feedback. Proceedings of the 37th International Conference on Neural Information Processing Systems. NIPS ’23. Curran Associates Inc, Red Hook, NY, USA, Yao S, Zhao J, Yu D, Du N, Shafran I, Narasimhan K, Cao Y (2023) ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations (ICLR) Yu T, Xu A, Akkiraju R (2024) defense of rag in the era of long-context language models. arXiv preprint arXiv:2409.01666 Zhu Y, Kellermann A, Gupta A, Li P, Fang R, Bindu R, Kang D (2024) Teams of LLM agents can exploit zero-day vulnerabilities. https://doi.org/10.48550/arXiv.2406.01637 Author information Authors and Affiliations Contributions Chaojie Wei and Yangyang Geng conceived the study and designed the methodology. Qilong Wu and Jing Huang implemented the framework and conducted the experiments. Yunfeng Wang and Qianqiong Wu contributed to the data collection and analysis. Qiang Wei supervised the project. All authors read and approved the final manuscript. Corresponding authors Ethics declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Competing interests The authors declare that they have no competing interests. Additional information Publisher's Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Appendices A prompts This appendix provides simplified versions of the system prompts used by the core agents in the PwnAgent framework. These prompts define the agents’ roles, CoT logic, tool usage specifications, and output formats. MainOrchestrator system prompt (Simplified) ProgramAnalyzer system prompt (Simplified) MeasurementExpert system prompt (Simplified) ExploitCrafter system prompt (Simplified) B Pwn benchmark overview Table 9 summarizes the pwn benchmark dataset used in the evaluation, with a dedicated field for source-platform provenance together with challenge types, file formats, enabled security measures, challenge difficulty, and intended exploit methods. Table 10 provides SHA-256 checksums for the exact binaries used in the evaluation. None of the 66 benchmark challenges was distributed with a challenge-specific libc; accordingly, no libc artifact was provided to the agents or included in the checksum table. Rights and permissions Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. About this article Cite this article Wei, C., Geng, Y., Wang, Y. et al. Pwnagent: a knowledge-guided multi-agent system for automatic exploit generation. Cybersecurity 9, 224 (2026). https://doi.org/10.1186/s42400-026-00649-5 Received: Accepted: Published: Version of record: DOI: https://doi.org/10.1186/s42400-026-00649-5 Keywords - Automatic exploit generation - Multi-agent systems - Large language models - Knowledge-guided reasoning - Dynamic introspection

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.