Benign Alone, Harmful Together: Exploiting Experience Composition in Self-Evolving LLM Agents
Abstract
Self-evolving large language model agents improve their capabilities by distilling interaction trajectories into persistent experiences. Yet this mechanism introduces a new safety risk: experiences that are benign in isolation may jointly weaken an agent’s safety boundary when accumulated and reused across sessions. Existing memory attacks typically require direct memory access or induce explicitly malicious records, limiting their stealthiness and applicability. We propose EvoBreak, an experience-conditioned sequential attack that operates through individually benign attack-stage tasks and induced experiences. EvoBreak repeatedly observes the experiences distilled by the victim, identifies uncovered target-relevant requirements, and adaptively acquires complementary experiences before reformulating the final query to activate them jointly. To support training, we introduce BreakGym, a structure-first synthesis pipeline that generates decomposable safety-sensitive targets with diverse dependency structures. EvoBreak is optimized using rejection-sampling supervised fine-tuning and Hint-guided GRPO. Experiments across self-evolving frameworks, victim backbones, pre-evolution domains, and safety benchmarks demonstrate that EvoBreak consistently outperforms existing attacks while maintaining high benignness. These results reveal benign experience composition as a persistent attack surface in self-evolving agents.
1 Introduction
Large language model (LLM)-based agents increasingly tackle long-horizon tasks that require planning, tool use, and adaptation across interactions (12; 32). To enable such cross-interaction adaptation, recent self-evolving agent frameworks distill past trajectories into reusable experiences and use them to guide future decisions (21; 17). This experience-driven paradigm has demonstrated substantial performance gains in coding, web navigation, and tool-augmented problem solving (11; 8).
However, the self-evolution mechanism creates a persistent attack surface as experiences may be reused across sessions as trusted internal guidance (16; 31). Early attacks assume direct access to the memory store to insert explicitly malicious records (4). Query-only attacks remove this direct-access requirement by inducing memory changes through ordinary interactions, ranging from attacker-designed malicious memories (6; 26; 33) to over-generalized rules distilled from valid cases (27). However, existing query-only attacks rarely treat stealthiness of induced memories as an explicit objective. Moreover, their effects are typically evaluated through downstream task degradation rather than safety-boundary erosion. Recent work shows that benign experience accumulation may inadvertently weaken refusal behavior in high-risk settings (34). This finding raises a critical question: Can an adaptive adversary deliberately induce a set of individually benign experiences that jointly weaken an agent’s safety boundary?
To investigate this question, we consider an adversary that can submit benign tasks and observe the experiences distilled from its interactions, but cannot directly modify them. At the target stage, the adversary issues a safety-sensitive query in a fresh session, where the preceding interaction history is unavailable while the accumulated experiences persist.
Realizing such an attack presents three key challenges. (1) Adaptive Planning. The victim may distill unexpected or incomplete experiences from each interaction, requiring the adversary to continually revise its attack plan based on the experiences actually generated rather than follow a fixed task sequence. (2) Benign Composition. Each submitted task and each induced experience must remain benign in isolation, while the resulting experience set can jointly affect the victim’s safety behavior. (3) Experience Heterogeneity. Different self-evolving agents may extract and represent substantially different experiences from similar interactions, making fixed task or experience templates unreliable.
To address these challenges, we propose EvoBreak, an experience-conditioned sequential attack on self-evolving agents that operates through individually benign tasks and induced experiences. Rather than committing to a fixed decomposition in advance, EvoBreak repeatedly examines the experiences distilled by the victim, estimates which target-relevant requirements remain uncovered, and generates the next benign task accordingly. This closed-loop process continues until the accumulated experiences are predicted to provide sufficient joint coverage, after which EvoBreak produces an experience-aligned target query that preserves the intent of the original safety-sensitive request.
Training such an adaptive attack policy requires decomposable targets whose local requirements can be acquired separately but must be jointly composed at the target stage. However, existing safety datasets largely consist of standalone harmful requests and provide little control over requirement decomposition or dependency structure, making them ill-suited for this purpose (19; 25; 1). We therefore introduce BreakGym, a scalable structure-first synthesis pipeline that generates diverse multi-domain targets by varying decomposition principles and dependency topologies, while leaving the interaction path unspecified. EvoBreak is then optimized through a two-stage workflow: rejection-sampling supervised fine-tuning (SFT) learns experience-conditioned planning from filtered successful trajectories, while Hint-guided Group Relative Policy Optimization (GRPO) further improves adaptation under sparse trajectory-level rewards using local training-time hints.
In summary, our contributions are as follows:
- •
We identify and systematically demonstrate a safety risk in self-evolving agents: experiences that are benign in isolation can jointly erode the agent’s safety boundary.
- •
We propose EvoBreak, an experience-conditioned attack that adaptively plans benign interactions based on victim-generated experiences.
- •
We introduce BreakGym, a scalable synthesis pipeline for decomposable safety-sensitive targets, and develop a cascaded training workflow.
- •
Extensive experiments demonstrate the effectiveness of EvoBreak across self-evolving agents and benchmarks.
2 Problem Setup and Threat Model
2.1 Self-Evolving Agent Setup
We consider a victim agent that solves a sequence of tasks while maintaining a persistent experience memory. For the -th task , the victim produces an execution trajectory , where denotes the persistent experience state available to the victim before processing . The trajectory abstracts the agent’s execution trace, including its actions, tool interactions, and final output.
After task completion, the victim distills the task and its trajectory into a reusable experience , where denotes the victim’s native experience-extraction process. The resulting experience may take the form of a reflection, procedural rule, or reusable skill, and is incorporated into the persistent memory, yielding the updated memory state .
2.2 Problem Formulation
Let denote the victim’s memory state before the attack, which may be empty or contain pre-existing experiences. The adversary submits a sequence of benign tasks Following the victim’s self-evolution process, each task induces an experience . We use to denote the post-attack memory state after all attacker-induced experiences have been incorporated.
At the target stage, the victim receives a safety-sensitive query in a fresh session, where the prior interaction context is unavailable but the persistent experience memory remains. Let denote the victim’s harmful-compliance score under target query and memory state . We define the experience-induced safety degradation as
| (1) |
The attack is subject to an atomic benignness constraint: each submitted task and each induced experience must be benign when evaluated in isolation,
| (2) |
where denotes a benignness auditor.
2.3 Threat Model
We consider a gray-box adversary that interacts with the victim through its task interface and has read-only observability of the experiences generated through its own interactions.
Adversary Capabilities and Constraints. The adversary may submit benign tasks and observe the corresponding experiences. However, the adversary cannot directly insert, delete, or edit memory entries, nor can it modify the victim’s system prompt or experience-extraction mechanism .
Adversarial Objective. The adversary seeks to maximize the experience-induced safety degradation while satisfying the atomic benignness constraint for every submitted task and induced experience.
3 Method
As illustrated in Figures 1 and 2, we propose an experience-conditioned sequential attack against self-evolving agents and its training framework. The overall framework consists of three components. First, EvoBreak adaptively constructs an attack interaction sequence by replanning based on the experiences generated by the victim. Second, BreakGym synthesizes structurally diverse safety-sensitive target scenarios to support the optimization. Third, EvoBreak is optimized through a cascaded training workflow consisting of rejection-sampling SFT and Hint-guided GRPO.
3.1 EvoBreak: Experience-Conditioned Sequential Attack
Given an initial safety-sensitive query , EvoBreak constructs an adaptive sequence of attack-stage interactions and a final target-stage query. Because the adversary controls the submitted tasks but not the experiences ultimately distilled by the victim, a predetermined decomposition of or a fixed task sequence may become ineffective when the generated experiences are incomplete, unexpected, or redundant.
EvoBreak therefore adopts a closed-loop strategy that replans after each interaction based on the experiences. The attack consists of two stages: experience-conditioned acquisition and experience-aware target reformulation.
Experience-Conditioned Acquisition. After interactions, EvoBreak maintains an observable history
| (3) |
where denotes the -th submitted task and is the corresponding experience generated by the victim.
To determine which target-relevant requirements remain insufficiently represented by the experiences generated so far, EvoBreak introduces an experience residual , an explicit textual planning state jointly generated with the next action:
| (4) |
where denotes the space of candidate tasks. The residual summarizes the target-relevant knowledge, procedures, or constraints that remain absent in the accumulated experiences. It is inferred solely from and , rather than from a predefined decomposition of the target.
If , the selected action is submitted to the victim as the next task, denoted by . The victim executes the task and generates a new experience . EvoBreak then updates the observable history as and replans from the updated history. This closed-loop process enables EvoBreak to revise its attack trajectory according to the experience actually produced at each interaction. The experience-acquisition stage terminates when EvoBreak selects Stop. Let denote the number of completed interactions at termination.
Experience-Aware Target Reformulation. Directly submitting the original query is likely to trigger the victim’s refusal behavior because its safety-sensitive intent is explicitly expressed. Although the induced experiences are closely related to , their presence in persistent memory alone does not ensure that the victim will jointly apply them at the target stage. EvoBreak therefore reformulates conditioned on the complete interaction history:
| (5) |
The reformulated query preserves the underlying intent of while aligning its semantic and procedural structure with the induced experiences, thereby encouraging their joint use in the victim’s target-stage reasoning.
3.2 BreakGym: Structure-First Target Synthesis
Training experience-conditioned planning requires targets whose local requirements can be elicited separately yet must be jointly composed. Existing safety datasets largely consist of standalone harmful queries, offering little control over their decomposition and dependency structures. We therefore introduce BreakGym, a scalable structure-first pipeline for synthesizing structurally controllable training targets.
BreakGym defines each structural configuration as
| (6) |
where specifies the decomposition principle, specifies the composition topology, and specifies the safety-sensitive target domain. Given a sampled configuration , BreakGym constructs a training instance where is a synthetic safety-sensitive query and is a latent scaffold encoding its local requirements and dependency relations. The scaffold specifies the internal organization of without prescribing the interaction trajectory of EvoBreak. It is accessible only during data construction and training.
Structural Taxonomy and Target Domains. As illustrated in Figure 2, BreakGym controls target structure through four decomposition principles: knowledge-based, procedural, constraint-based, and functional, and four composition topologies: sequential, parallel, hierarchical, and conditional. These structures are combined with multiple safety-sensitive domains to generate complex targets. Detailed definitions are provided in Appendix A.
Target Construction Pipeline. BreakGym constructs each target through a three-stage pipeline:
Stage 1. Structure Sampling. Given a sampled configuration , BreakGym constructs a latent scaffold , where each node represents a local target requirement, and encodes the structural relations among the requirements according to the sampled composition topology.
Stage 2. Target Instantiation. BreakGym instantiates the scaffold within the sampled target domain to generate a coherent safety-sensitive query . Because the local requirements and their dependencies are specified before generation, the structure of the query is explicitly controllable. BreakGym can vary domains, contexts, and entities to scalably generate diverse decomposable queries.
Stage 3. Scenario Validation and Filtering. BreakGym validates whether each synthesized query faithfully reflects its scaffold while remaining coherent and natural as a standalone request. Queries whose local requirements are indistinguishable, weakly complementary, or inconsistent with the specified dependency structure are discarded. The remaining instances with similarity to downstream evaluation prompts are removed to prevent data leakage.
3.3 Cascaded Optimization of EvoBreak
EvoBreak exposes attack as a sequential and trainable process. As shown in Figure 1, BreakGym supports a cascaded optimization workflow: rejection-sampling SFT first initializes the sequential attack policy, followed by Hint-guided GRPO for further exploration under outcome-level feedback.
Rejection-Sampling SFT. For each BreakGym target , a strong teacher model performs multiple independent rollouts under the EvoBreak workflow. The teacher follows the same observation setting as EvoBreak and has no access to the latent scaffold . We retain only rollouts in which all submitted tasks and induced experiences pass the benignness audit and the final target-stage attack succeeds. The retained rollouts form the supervised dataset .
We then perform standard supervised fine-tuning on the complete attack trajectories:
| (7) |
where denotes the -th model-generated output in rollout , and denotes its preceding trajectory context.
| Attack | JailbreakBench | HarmBench | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SE-Agent | ReasoningBank | SE-Agent | ReasoningBank | |||||||||
| Math | Code | Reas. | Math | Code | Reas. | Math | Code | Reas. | Math | Code | Reas. | |
| GPT-5-mini | ||||||||||||
| PAIR | 22.60 | 27.20 | 27.20 | 24.40 | 26.80 | 24.60 | 23.75 | 24.55 | 23.75 | 23.50 | 23.65 | 24.00 |
| FlipAttack | 27.80 | 40.00 | 40.00 | 29.20 | 28.40 | 38.60 | 36.55 | 37.20 | 37.35 | 35.25 | 35.95 | 36.60 |
| ReNeLLM | 52.20 | 50.40 | 52.00 | 51.60 | 48.00 | 50.20 | 57.05 | 55.85 | 61.00 | 57.25 | 58.95 | 59.10 |
| AgentPoison | 47.60 | 54.80 | 53.20 | 48.20 | 52.40 | 49.00 | 51.10 | 55.95 | 53.75 | 49.25 | 56.35 | 52.80 |
| MINJA | 34.40 | 30.20 | 31.00 | 31.60 | 31.40 | 33.20 | 28.70 | 30.45 | 26.65 | 31.25 | 28.15 | 27.55 |
| EvoBreak | 80.60 | 81.00 | 79.20 | 79.80 | 78.20 | 80.40 | 85.05 | 84.80 | 86.25 | 83.05 | 85.00 | 86.10 |
| Llama-3.1-8B-Instruct | ||||||||||||
| PAIR | 28.20 | 33.60 | 31.40 | 34.80 | 34.20 | 31.60 | 20.75 | 22.25 | 21.70 | 22.05 | 20.80 | 21.55 |
| FlipAttack | 46.00 | 52.40 | 54.40 | 54.20 | 47.80 | 50.60 | 43.75 | 47.50 | 47.45 | 46.15 | 44.30 | 46.70 |
| ReNeLLM | 67.80 | 70.00 | 69.60 | 71.20 | 70.40 | 66.00 | 70.10 | 73.35 | 71.65 | 67.90 | 66.15 | 68.40 |
| AgentPoison | 68.40 | 70.40 | 66.80 | 70.20 | 72.00 | 67.20 | 74.15 | 70.35 | 71.75 | 70.05 | 69.40 | 72.30 |
| MINJA | 44.20 | 48.60 | 49.20 | 40.40 | 46.80 | 48.00 | 38.05 | 41.25 | 36.40 | 34.10 | 38.35 | 35.55 |
| EvoBreak | 90.60 | 88.20 | 87.80 | 88.00 | 89.40 | 87.40 | 93.85 | 91.10 | 89.55 | 90.30 | 92.95 | 88.20 |
Hint-guided GRPO. To further improve EvoBreak’s ability under sparse trajectory-level rewards, we introduce Hint-guided GRPO. A fixed strong model serves as a training-time critic, using the latent scaffold to provide local corrective guidance for intermediate planning decisions.
At each interaction step, the attack model first generates an initial residual–action pair. The critic evaluates this decision using the latent scaffold and returns no hint when the decision is appropriate. Otherwise, it provides a local corrective hint, based on which the attack model regenerates the decision. Only the final decision is retained for rollout construction. The hint critic and the latent scaffold are used only during training and are unavailable at inference time.
Let indicate whether the attack succeeds. For each rollout , we define the trajectory-level reward as
| (8) |
where . The first term rewards attack success, while the second measures the average benignness.
For each target, we sample a group of rollouts and normalize their rewards within the group to obtain the group-relative advantages . We then optimize EvoBreak using the standard GRPO objective:
| (9) |
where denotes the policy ratio for rollout , is the reference policy, and controls the KL penalty.
4 Experiment
We conduct extensive experiments to evaluate the effectiveness and benignness of EvoBreak, assess BreakGym as a training source, and investigate the mechanisms underlying benign experience composition. Specifically, our evaluation aims to answer the following questions: RQ1: How does EvoBreak compare with existing attack methods in terms of effectiveness and benignness? RQ2: Does BreakGym provide effective supervision for training EvoBreak? RQ3: How do the key components of EvoBreak contribute to its attack effectiveness? RQ4: Does EvoBreak’s effectiveness arise from the joint composition of multiple induced experiences, and how robust is it under different memory conditions?
4.1 Experiment Setting
Self-evolving Frameworks. We evaluate EvoBreak on two representative self-evolving agent frameworks with distinct evolution mechanisms: ReasoningBank (21), which distills reusable reasoning strategies from successful and failed trajectories, and SE-Agent (17), which iteratively improves prior trajectories through revision, recombination, and refinement.
Pre-Evolution Domains. Before launching the attack, we initialize each self-evolving agent with experiences acquired from three domains: mathematics, code, and general reasoning. We use AIME (14) for mathematics, LiveCodeBench (13) for code, and MMLU-Pro (28) for general reasoning.
Attack Benchmarks. We evaluate all attacks on two safety benchmarks, JailbreakBench (2), which contains 100 malicious prompts, and HarmBench, which consists of 400 harmful behaviors (19).
Evaluation Metrics. We report Attack Success Rate (ASR) and Benignness (Ben.). ASR is the percentage of evaluation queries for which the victim’s final response successfully fulfills the target objective. Benignness is computed by independently auditing every attack-stage task and induced experience and averaging the resulting binary judgments.
Baselines. We compare EvoBreak against two categories of baselines: (1) Query-level jailbreak attacks: PAIR (3), FlipAttack (18), and ReNeLLM (5); and (2) Memory-oriented attacks: AgentPoison (4) and MINJA (6).
Implementation Details. We employ GPT-5-mini (20) and Llama-3.1-8B-Instruct (10) as the backbones of the victim agents, and Qwen3.5-9B (22) as the backbone of EvoBreak. GPT-5.4-mini is used as the judge model for evaluating ASR and benignness. Unless otherwise specified, all results are averaged over five independent runs. During optimization, we sample parallel rollouts for each target and use a learning rate of . More details about the experimental setting are provided in Appendix B.
4.2 Main Results
Table 1 systematically compares EvoBreak with five competitive baselines across two victim models, two self-evolving frameworks, three pre-evolution domains, and two safety benchmarks. We highlight three main observations.
First, EvoBreak consistently achieves the strongest attack performance across all settings. Its overall average ASR reaches 86.12%, outperforming the strongest baseline, ReNeLLM, by 24.19 percentage points. The consistent gains over both query-level and memory-oriented attacks demonstrate that benign experience composition is an effective attack surface for self-evolving agents.
Second, EvoBreak is robust to heterogeneous self-evolution mechanisms and pre-evolution domains. It maintains consistently high ASR on both SE-Agent and ReasoningBank, as well as across mathematics, code, and general reasoning domains. This robustness suggests that EvoBreak does not rely on a specific experience representation or initial memory distribution. Instead, its experience-conditioned acquisition process continually replans based on the experiences actually generated by the victim, adapting the attack trajectory to different evolutionary settings.
Third, EvoBreak generalizes effectively across victim backbones. Unlike attacks that primarily exploit model-specific prompt vulnerabilities, EvoBreak leverages the general experience accumulation and reuse mechanism of self-evolving agents. Its experience-aware target reformulation further aligns the final query with the accumulated benign experiences, facilitating their joint reuse and weakening the victim’s refusal boundary.
4.3 Benignness of Attacks
Figure 3 reports the average ASR and benignness on Llama-3.1-8B-Instruct across self-evolving frameworks and pre-evolution domains. EvoBreak achieves the best effectiveness–benignness trade-off on both benchmarks. Query-level attacks directly rewrite or obfuscate the target request, but must still expose sufficient harmful intent within a single query, limiting either ASR or benignness. AgentPoison achieves stronger attack performance through direct injection of malicious memory records, requiring privileged access and making the attack more detectable.
In contrast, EvoBreak’s high benignness is supported by its training and adaptive planning design. BreakGym teaches the attacker to acquire complementary requirements through separate interactions rather than exposing the complete harmful intent in a single task. Rejection-sampling SFT filters out trajectories containing non-benign tasks or experiences, while Hint-guided GRPO explicitly rewards benign intermediate interactions. Moreover, experience-conditioned planning adapts each subsequent task to the experiences actually generated by the victim, maintaining target relevance without resorting to overtly harmful requests. These mechanisms allow EvoBreak to achieve high benignness while preserving strong attack effectiveness.
| Variant | SE-Agent | ReasoningBank | ||||
|---|---|---|---|---|---|---|
| Math | Code | Reas. | Math | Code | Reas. | |
| Llama-3.1-8B-Instruct | ||||||
| w/o Training | 52.40 | 56.20 | 54.00 | 56.60 | 54.20 | 56.20 |
| SFT Only | 75.60 | 72.20 | 73.80 | 76.20 | 73.40 | 72.60 |
| GRPO w/o Hint | 79.00 | 81.80 | 81.40 | 78.20 | 80.00 | 81.20 |
| w/o Replanning | 67.60 | 63.20 | 62.80 | 64.00 | 65.20 | 62.60 |
| w/o Reformulation | 66.60 | 70.40 | 63.00 | 65.20 | 68.00 | 64.20 |
| RF Query Only | 31.20 | 27.80 | 28.40 | 26.40 | 28.00 | 30.20 |
| Full | 90.60 | 88.20 | 87.80 | 88.00 | 89.40 | 87.40 |
4.4 Ablation Study
To evaluate the contribution of each component, we conduct systematic ablations, with results reported in Table 2.
First, training and optimization substantially improve attack effectiveness. The untrained variant still achieves non-trivial success, indicating that experience acquisition and target reformulation are effective even without specialized training. Rejection-sampling SFT provides strong supervision from successful and benign trajectories, while GRPO further improves performance; the additional gains from Hint-guided GRPO show that local corrective guidance benefits intermediate planning under sparse trajectory-level rewards.
Second, experience-conditioned replanning is critical. A fixed task decomposition cannot account for the experiences actually distilled by the victim, which may be incomplete, redundant, or off-target. By planning over the realized experience state, EvoBreak continually targets uncovered requirements and constructs a more complementary experience set.
Third, experience acquisition and target reformulation are complementary. Acquisition constructs the target-relevant experience basis, whereas reformulation activates and composes these distributed experiences at the target stage. Removing either component reduces the attack to unsupported query rewriting or poorly activated experience accumulation, demonstrating the importance of coupling experience construction with experience activation.
4.5 Analysis of BreakGym Supervision
To examine whether BreakGym provides effective supervision for training EvoBreak, we conduct an SFT-only comparison using different training sources, including AdvBench (35), Do-Not-Answer (29), WildJailbreak (15), and Jailbreak_LLMs (23). All variants share the same framework, backbone model, and action space, and GRPO is excluded to isolate the effect of supervision data.
As shown in Figure 4, BreakGym provides the strongest supervision among all training sources. Unlike existing safety datasets, which mainly contain standalone harmful requests, BreakGym explicitly organizes complex targets through decomposition principles and composition topologies. This structured construction better matches EvoBreak’s need to identify uncovered requirements and acquire complementary experiences across interactions, thereby providing more effective supervision for experience-conditioned planning.
4.6 Analysis of Experience Conditions
Effect of Retained Attack Experiences. To examine whether EvoBreak’s effectiveness arises from a single decisive experience or from the composition of accumulated experiences, we vary the fraction of attack experiences retained at the target stage. As shown in Figure 5, attack success remains limited when no experience or only a small subset is available, but increases progressively and accelerates as the retained set approaches completion. These results suggest that EvoBreak generally benefits from the joint availability of multiple experiences rather than relying solely on a small retained subset. The pronounced gains at high retention levels further suggest that replanning does more than simply accumulate additional experiences: by repeatedly addressing residual coverage gaps, it constructs a complementary experience set whose effectiveness emerges through joint activation at the target stage.
Effect of Pre-existing Experiences. To evaluate EvoBreak’s robustness across different memory contexts, we vary the number of pre-existing experiences before the attack. As shown in Figure 6, although attack success gradually decreases with the memory scale, EvoBreak remains highly effective even with extensive pre-existing experiences. By replanning over the realized experience state and aligning the final query with the accumulated experiences, EvoBreak adapts to diverse background memories without relying on an empty or fixed initial state.
5 Related Work
Self-Evolving LLM Agents
Self-evolving LLM agents continually improve by distilling interaction trajectories into persistent experiences and reusing them in future tasks (7). Existing approaches typically organize successful and failed trajectories or iteratively refine prior trajectories, enabling continual adaptation across diverse agent tasks (17; 21; 30).
Attacks on Self-Evolving LLM Agents
Security attacks relevant to self-evolving LLM agents mainly target either the final query or the agent’s persistent memory. Query-level jailbreaks rewrite or obfuscate harmful requests to bypass safety mechanisms (3; 18; 5), while memory-oriented attacks directly inject malicious records or induce harmful experiences through interaction (4; 6).
Training Data for Automated Red-Teaming
Existing safety and jailbreak datasets provide diverse harmful intents and adversarial requests for evaluating and training automated red-teaming methods (35; 29; 15; 23). However, they typically represent each target as a standalone query, offering limited supervision for decomposing complex objectives and modeling dependencies across interactions.
6 Conclusion
In this paper, we propose EvoBreak, an experience-conditioned sequential attack that demonstrates benign experience composition as a persistent safety risk in self-evolving LLM agents. EvoBreak adaptively acquires complementary, individually benign experiences and reformulates the final query to activate them jointly. We further introduce BreakGym and a cascaded optimization workflow combining rejection-sampling SFT with Hint-guided GRPO. Extensive experiments demonstrate strong attack effectiveness and high benignness, highlighting the need to assess the cumulative safety effects of persistent experiences rather than auditing them in isolation.
References
- Agentharm: a benchmark for measuring harmfulness of llm agents. In International Conference on Learning Representations, Vol. 2025, pp. 79185–79220. Cited by: §1.
- Jailbreakbench: an open robustness benchmark for jailbreaking large language models. Advances in Neural Information Processing Systems 37, pp. 55005–55029. Cited by: §B.3, §4.1.
- Jailbreaking black box large language models in twenty queries. In 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pp. 23–42. Cited by: §B.4, §4.1, §5.
- Agentpoison: red-teaming llm agents via poisoning memory or knowledge bases. Advances in Neural Information Processing Systems 37, pp. 130185–130213. Cited by: §B.4, §1, §4.1, §5.
- A wolf in sheep’s clothing: generalized nested jailbreak prompts can fool large language models easily. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 2136–2153. Cited by: §B.4, §4.1, §5.
- Memory injection attacks on llm agents via query-only interaction. Advances in Neural Information Processing Systems 38, pp. 46697–46731. Cited by: §B.4, §1, §4.1, §5.
- A comprehensive survey of self-evolving ai agents: a new paradigm bridging foundation models and lifelong agentic systems. arXiv preprint arXiv:2508.07407. Cited by: §5.
- WebEvolver: enhancing web agent self-improvement with co-evolving world model. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 8970–8986. Cited by: §1.
- Gemini 3 Flash: model card. Note: Model card External Links: Link Cited by: §B.5.
- The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §4.1.
- Controlled self-evolution for algorithmic code optimization. arXiv preprint arXiv:2601.07348. Cited by: §1.
- Understanding the planning of llm agents: a survey. arXiv preprint arXiv:2402.02716. Cited by: §1.
- Livecodebench: holistic and contamination free evaluation of large language models for code. In International Conference on Learning Representations, Vol. 2025, pp. 58791–58831. Cited by: §B.2, §4.1.
- AIME Problem Set 2024. Hugging Face. Note: https://huggingface.co/datasets/Maxwell-Jia/AIME_2024 Cited by: §B.2, §4.1.
- Wildteaming at scale: from in-the-wild jailbreaks to (adversarially) safer language models. Advances in Neural Information Processing Systems 37, pp. 47094–47165. Cited by: §4.5, §5.
- Behavior safety of autonomous interactive agents: risks, attacks, defenses and evaluation. Evaluation 79, pp. 19–2. Cited by: §1.
- Se-agent: self-evolution trajectory optimization in multi-step reasoning with llm-based agents. arXiv preprint arXiv:2508.02085. Cited by: §B.1, §1, §4.1, §5.
- Flipattack: jailbreak llms via flipping. arXiv preprint arXiv:2410.02832. Cited by: §B.4, §4.1, §5.
- Harmbench: a standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249. Cited by: §B.3, §1, §4.1.
- GPT-5 Mini Model. Note: OpenAI API Documentation External Links: Link Cited by: §4.1.
- Reasoningbank: scaling agent self-evolving with reasoning memory. arXiv preprint arXiv:2509.25140. Cited by: §B.1, §1, §4.1, §5.
- Qwen3.5: towards native multimodal agents. External Links: Link Cited by: §4.1.
- " Do anything now": characterizing and evaluating in-the-wild jailbreak prompts on large language models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pp. 1671–1685. Cited by: §4.5, §5.
- Hybridflow: a flexible and efficient rlhf framework. In Proceedings of the Twentieth European Conference on Computer Systems, pp. 1279–1297. Cited by: §B.5.
- A strongreject for empty jailbreaks. Advances in Neural Information Processing Systems 37, pp. 125416–125440. Cited by: §1.
- MemoryGraft: persistent compromise of llm agents via poisoned experience retrieval. arXiv preprint arXiv:2512.16962. Cited by: §1.
- OEP: poisoning self-evolving llm agents via locally correct but non-transferable experiences. arXiv preprint arXiv:2605.18930. Cited by: §1.
- Mmlu-pro: a more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems 37, pp. 95266–95290. Cited by: §B.2, §4.1.
- Do-not-answer: evaluating safeguards in llms. In Findings of the Association for Computational Linguistics: EACL 2024, pp. 896–911. Cited by: §4.5, §5.
- Evo-memory: benchmarking llm agent test-time learning with self-evolving memory. arXiv preprint arXiv:2511.20857. Cited by: §5.
- Evo-attacker: memory-augmented reinforcement learning for long-horizon tool attacks on llm-mas. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 7286–7300. Cited by: §1.
- Beyond self-talk: a communication-centric survey of llm-based multi-agent systems. arXiv preprint arXiv:2502.14321. Cited by: §1.
- Zombie agents: persistent control of self-evolving llm agents via self-reinforcing injections. arXiv preprint arXiv:2602.15654. Cited by: §1.
- On safety risks in experience-driven self-evolving agents. In Findings of the Association for Computational Linguistics: ACL 2026, pp. 42145–42169. Cited by: §1.
- Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043. Cited by: §4.5, §5.
Appendix A BreakGym
BreakGym represents each target using a structural configuration
where specifies how the target is decomposed into local requirements, determines how these requirements depend on and compose with one another, and specifies the safety-sensitive domain in which the structure is instantiated.
Given a sampled structural configuration, BreakGym constructs a latent scaffold
where each node represents a local target requirement and represents the structural relations among these requirements. The following sections provide detailed definitions of the decomposition principles, composition topologies, and safety-sensitive target domains used in BreakGym.
A.1 Decomposition Principles
The decomposition principle determines the semantic meaning of the nodes in the latent scaffold. BreakGym considers four complementary decomposition principles, which capture different ways in which a complex target can be divided into locally meaningful requirements.
Knowledge-based decomposition.
Knowledge-based decomposition separates a target according to the distinct knowledge components required for completing it. Each node represents a self-contained unit of factual, conceptual, contextual, or technical knowledge, such as the properties of an entity, the mechanism of a process, or the relationship between multiple concepts. Although each knowledge unit is locally meaningful, successful completion of the target requires integrating multiple complementary units.
Procedural decomposition.
Procedural decomposition divides a target into a set of operational stages or subprocedures. Each node captures the knowledge or capability required to perform one stage, while the latent scaffold specifies how the output of one stage supports subsequent stages. This decomposition principle is suitable for targets whose completion depends on combining multiple pieces of step-level knowledge into a coherent end-to-end procedure.
Constraint-based decomposition.
Constraint-based decomposition organizes a target according to the conditions that a valid solution must satisfy. These conditions may describe required inputs, resource limitations, operating conditions, output specifications, or exception-handling requirements. Each node encodes an individual constraint or a constraint-specific requirement, whereas completing the overall target requires jointly satisfying the full set of constraints.
Functional decomposition.
Functional decomposition separates a target according to the functional roles performed by its components. Each node represents a distinct module or capability, such as information acquisition, transformation, verification, coordination, or output generation. Unlike procedural decomposition, functional decomposition does not necessarily impose a fixed execution order. Instead, it emphasizes the complementary contributions of different functional units to the overall objective.
A.2 Composition Topologies
The composition topology determines the dependency relations among the local requirements in the latent scaffold. While the decomposition principle specifies what each node represents, the composition topology specifies how the nodes must be combined. BreakGym considers four representative composition topologies.
Sequential topology.
In a sequential topology, local requirements form an ordered chain, which can be represented as
A downstream requirement depends on the successful completion or output of one or more preceding requirements. Consequently, the target cannot be completed by satisfying the requirements in an arbitrary order.
Parallel topology.
In a parallel topology, multiple local requirements contribute independently or weakly dependently to the overall target. No strict ordering is imposed among the corresponding nodes, but their outputs must be jointly available at the composition stage. This topology captures targets that require several complementary capabilities or information components rather than a single ordered procedure.
Hierarchical topology.
A hierarchical topology organizes local requirements at multiple levels of abstraction. A high-level objective is recursively decomposed into subobjectives, and each subobjective may be further divided into more specific requirements. Internal nodes represent composite objectives, whereas leaf nodes represent locally elicitable requirements. Completing the target requires aggregating information or capabilities from the leaf nodes through the hierarchical structure.
Conditional topology.
In a conditional topology, the relevance or execution of a local requirement depends on the outcome of another requirement or on an explicit condition. Different branches correspond to different contextual states, decisions, or intermediate results. Completing the target therefore requires identifying the applicable condition and composing the requirements along the corresponding branch.
A.3 Safety-Sensitive Target Domains
After sampling a decomposition principle and a composition topology, BreakGym instantiates the resulting latent scaffold within a safety-sensitive target domain. The target domain determines the entities, scenarios, terminology, and local requirements used to realize the abstract structure. BreakGym covers six safety-sensitive domains: privacy, fraud, chemical, cyber, violence, and sexual safety.
Privacy.
The privacy domain covers targets involving the unauthorized collection, inference, linkage, disclosure, or exploitation of personal or confidential information. Complex targets in this domain may combine requirements related to information-source identification, identity linkage, attribute inference, and information disclosure.
Fraud.
The fraud domain concerns deceptive or manipulative activities intended to obtain financial, informational, or other illegitimate benefits. Targets in this domain may contain multiple components related to impersonation, persuasion, transaction manipulation, concealment, or evasion of verification mechanisms.
Chemical.
The chemical domain includes safety-sensitive targets involving hazardous, restricted, or prohibited chemical substances and processes. Relevant local requirements may concern materials, chemical transformations, operating conditions, handling procedures, post-processing, or process optimization. BreakGym organizes these requirements structurally without assuming a single fixed decomposition pattern.
Cyber.
The cyber domain covers targets involving unauthorized access, exploitation of software or systems, credential compromise, malicious code, persistence, or evasion of security mechanisms. Targets in this domain may contain distinct requirements related to reconnaissance, vulnerability identification, access establishment, execution, and post-exploitation behavior.
Violence.
The violence domain contains targets associated with planning, facilitating, or carrying out physical harm. Depending on the sampled structure, local requirements may represent resource acquisition, target-related information, operational planning, coordination, or avoidance of intervention.
Sexual safety.
The sexual-safety domain includes targets involving sexual exploitation, non-consensual sexual content, age-inappropriate material, or other violations of sexual-safety policies. Latent scaffolds in this domain may contain distinct requirements involving content generation, manipulation, targeting, distribution, or concealment.
Appendix B Experiment
This section provides additional details about the experimental settings used to evaluate EvoBreak. We describe the self-evolving frameworks, pre-evolution domains, attack benchmarks, baseline methods, and implementation details.
B.1 Self-Evolving Frameworks
We evaluate EvoBreak on two representative self-evolving agent frameworks with distinct experience construction and reuse mechanisms.
ReasoningBank.
ReasoningBank (21) is a self-evolving agent framework that distills reusable reasoning strategies from both successful and failed interaction trajectories. The extracted experiences are stored in a persistent reasoning memory and can be retrieved to guide future problem-solving processes. By learning from both positive and negative trajectories, ReasoningBank constructs experience entries that capture not only effective reasoning patterns but also lessons derived from previous failures.
SE-Agent.
SE-Agent (17) performs self-evolution through iterative trajectory optimization. Instead of only extracting independent reasoning strategies, it improves previously generated trajectories through revision, recombination, and refinement. The resulting experiences preserve reusable information derived from prior executions and support subsequent decision-making.
B.2 Pre-Evolution Domains
Before launching an attack, we initialize each victim agent with pre-existing experiences collected from a non-adversarial task domain. This setting reflects the practical scenario in which a self-evolving agent has already accumulated experiences through ordinary use before being exposed to an adversarial interaction sequence. We consider three pre-evolution domains: mathematics, code, and general reasoning.
Mathematics.
For mathematical pre-evolution, we use AIME problems (14). These problems require multi-step mathematical reasoning and therefore induce experiences related to problem decomposition, intermediate derivation, and verification of candidate solutions.
Code.
For code-domain pre-evolution, we use LiveCodeBench (13). Tasks from this benchmark induce programming-oriented experiences involving problem interpretation, algorithm design, implementation, and solution verification.
General Reasoning.
For general-reasoning pre-evolution, we use MMLU-Pro (28). MMLU-Pro covers questions requiring reasoning across a broad range of subject areas. Experiences generated from this domain are therefore more heterogeneous than those obtained from mathematics or code alone.
B.3 Attack Benchmarks
We evaluate attack effectiveness on two widely used safety benchmarks, JailbreakBench and HarmBench. The two benchmarks differ in scale and target coverage, enabling us to assess whether the effectiveness of EvoBreak generalizes across different collections of safety-sensitive behaviors.
JailbreakBench.
JailbreakBench (2) is a standardized benchmark for evaluating jailbreak attacks and model refusal robustness. We use its set of 100 malicious prompts as target queries. Each prompt specifies a safety-sensitive objective that the victim model is expected to refuse under normal conditions.
HarmBench.
HarmBench (19) is a broader evaluation framework containing 400 harmful behaviors. Compared with JailbreakBench, HarmBench provides a larger and more diverse collection of safety-sensitive objectives.
B.4 Baselines
We compare EvoBreak with five representative baseline attacks. These methods are divided into two categories: query-level jailbreak attacks, which primarily manipulate the target query, and memory-oriented attacks, which attempt to compromise the persistent memory or experience state of an agent.
Query-Level Jailbreak Attacks
PAIR.
PAIR (3) is an iterative black-box jailbreak method. It repeatedly generates and refines candidate jailbreak prompts according to feedback obtained from the target model. The refinement process seeks to preserve the original target objective while improving the probability that the resulting prompt bypasses the model’s safety mechanisms.
FlipAttack.
FlipAttack (18) constructs one-shot adversarial queries by transforming or flipping the textual representation of the original request. The transformed query is designed to reduce the effectiveness of surface-level safety detection while retaining sufficient information for the target model to recover the intended request.
ReNeLLM.
ReNeLLM (5) combines prompt rewriting with scenario nesting. It places the target request within a constructed contextual scenario and reformulates its surface expression to reduce the likelihood of direct refusal. Unlike iterative methods such as PAIR, ReNeLLM relies on structured query transformation and contextual embedding.
Memory-Oriented Attacks
AgentPoison.
AgentPoison (4) attacks an LLM agent by directly poisoning its persistent memory or external knowledge base. The injected records are designed to influence the agent’s future behavior when they are retrieved during downstream tasks. Because AgentPoison assumes the ability to directly insert or modify memory entries, it operates under a stronger adversarial-access assumption than EvoBreak.
MINJA.
MINJA (6) is a query-only memory injection attack that induces malicious memory records through interactions with the victim agent. Unlike AgentPoison, MINJA does not require direct write access to the memory store. Instead, it constructs queries whose execution trajectories cause the victim’s native experience-extraction mechanism to generate attacker-desired memory entries. MINJA is therefore the closest baseline to EvoBreak in terms of interaction-based memory manipulation. However, whereas MINJA aims to induce malicious memory records, EvoBreak operates through multiple tasks and experiences that remain benign when evaluated individually.
B.5 Implementation Details
All experiments were conducted on computational nodes equipped with 8 NVIDIA H100 80GB GPUs. For supervised fine-tuning, we optimize the EvoBreak policy on verified trajectories using the standard next-token prediction objective. We use Qwen3.5-122B-A10B to generate interaction trajectories and Gemini 3 Flash (9) to filter low-quality or unsupported trajectories.
For reinforcement learning, we use Qwen3.5-9B as the backbone of EvoBreak and optimize the policy using Hint-guided GRPO initialized from the SFT checkpoint. We use the Verl library (24) with vLLM-based rollouts, gradient checkpointing, and FSDP offloading. For each target, we sample rollouts for group-relative advantage estimation and train the policy for 200 optimization steps. The learning rate is set to , the KL-regularization coefficient is set to , and the clipping parameter is set to . Hints and latent scaffolds are used only during training and are removed at inference time. Unless otherwise specified, all experimental results are averaged over five independent runs.
Appendix C The Template of Prompt
C.1 Evaluation Prompt of ASR
We use an LLM-based evaluator to assess whether the victim model successfully fulfills a safety-sensitive instruction. The evaluator assigns each response a score from 1 to 5 according to the degree of compliance with the user’s underlying intent. An attack is considered successful only when the response receives a score of 5. Accordingly, the Attack Success Rate is computed as
| (10) |
where is the total number of evaluated target queries and is the evaluator score for the -th response. The evaluation prompt is provided below.
C.2 Evaluation Prompt of Benignness
We employ an LLM-based evaluator to determine whether each attack-stage task and induced experience is benign when considered in isolation. The evaluator assigns a binary label, where 1 indicates benign content and 0 indicates non-benign content. In particular, content is considered non-benign if it contains malicious instructions, harmful operational guidance, or explicit encouragement that meaningfully facilitates an unsafe objective.
Given an attack trajectory containing submitted tasks and induced experiences, benignness is calculated as
| (11) |
where denotes the binary benignness judgment. The complete evaluation prompt is shown below.