CausalDS: Benchmarking Causal Reasoning in Data-Science Agents Paper • 2607.08093 • Published 16 days ago • 5
view article Article ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration ibm-research • 24 days ago • 24
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 95 • 3
SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning Paper • 2602.19455 • Published Feb 23 • 1
Enterprise Agents and Benchmarks Collection Enterprise agent ecosystem featuring AssetOpsBench (industrial) and ITBench (SRE, FinOps, CISO), CUGA to accelerate AI Automation • 21 items • Updated about 1 month ago • 18
Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows Paper • 2605.24219 • Published May 26 • 9
Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents Paper • 2606.12674 • Published Jun 10 • 5
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Paper • 2606.19704 • Published Jun 18 • 41
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Paper • 2606.19704 • Published Jun 18 • 41
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Paper • 2606.19704 • Published Jun 18 • 41
Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents Paper • 2606.12674 • Published Jun 10 • 5
view reply Appreciate the nice writeup. Can we add a) Leaderboard, b) Benchmark https://github.com/IBM/AssetOpsBench
view article Article Beyond LLMs: Why Scalable Enterprise AI Adoption Depends on Agent Logic ibm-research • Jun 1 • 89