Abstract
Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not always design training and test tasks such that test-time gains can be attributed to training experience, and remain vulnerable to data contamination. We present GDPevo, an evolution-native benchmark grounded in GDP-related enterprise workflows, together with the fully automated data pipeline that generates it. Its core mechanism, rule hybridization, decomposes each enterprise workflow into atomic business rules, distributes subsets of these rules across training tasks, and recombines them in held-out test tasks so that test-time gains are attributable. GDPevo spans CRM, ERP, finance, healthcare, legal, and data-centric workflows. Its V1 release contains 120 tasks in 12 groups, with five training and five held-out test tasks per group. Full automation enables the pipeline to expand the suite to 240 tasks in 24 groups (V2) within two days, providing a practical response to contamination. Using GDPevo, we evaluate four agents, each comprising a harness and a model, under four supervision types. Self-evolution consistently improves held-out accuracy by up to 16.44 percentage points. But the best evolved agents remain far below the fully informed oracle ceiling of 91.6%, indicating that the self-evolution ability of current agents remains far from fully realized. We publicly release the pipeline, benchmark, and full evaluation results at https://github.com/Prism-Shadow/GDPevo.
Community
We release GDPevo, the first benchmark for evaluating agent self-evolution on GDP-related enterprise tasks.
We also open-source the fully automated pipeline that generates evolution-native benchmark data, providing a practical response to contamination.
We found the best evolved agents remain far below a fully informed oracle ceiling of 91.6%, indicating that the self-evolution ability of current agents remains far from fully realized.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer (2026)
- SEAGym: An Evaluation Environment for Self-Evolving LLM Agents (2026)
- EvoClawBench: Can Agents Learn Reusable Skills from Their Own Runs? (2026)
- Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction (2026)
- ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End? (2026)
- ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents (2026)
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.03764 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
