FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 5 days ago • 140
Meta$^n$: Recursive Self-Improvement through Emergent Depth Paper • 2608.24735 • Published 5 days ago • 12
AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI Paper • 2511.20686 • Published Nov 20, 2025
Meta^n: Recursive Self-Improvement through Emergent Depth Paper • 2608.24735 • Published 5 days ago • 12
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Paper • 2608.20169 • Published 6 days ago • 11
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 6 days ago • 200
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 10 days ago • 64
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published 11 days ago • 51
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve Paper • 2608.16884 • Published 13 days ago • 18
On-Policy Delta Distillation for Multilingual Math Reasoning Paper • 2608.05802 • Published 24 days ago • 32
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Paper • 2608.06301 • Published 24 days ago • 35
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published 27 days ago • 183
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 25 days ago • 23
Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation Paper • 2607.27372 • Published Jul 29 • 19