SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback Paper • 2608.13120 • Published 11 days ago • 30
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization Paper • 2608.20281 • Published 4 days ago • 10
MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use Paper • 2608.20202 • Published 4 days ago • 31
The Problem Is the Problem: Towards Scalable Mathematical Discovery Paper • 2608.16977 • Published 7 days ago • 4
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Paper • 2608.13558 • Published 11 days ago • 88
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Paper • 2608.14905 • Published 10 days ago • 29
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 9 days ago • 436
Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents Paper • 2608.15008 • Published 9 days ago • 15
Personalized Auto-Research: Towards a True AI Co-Scientist Paper • 2608.14881 • Published 10 days ago • 5
Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection Paper • 2608.16393 • Published 6 days ago • 6
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published 6 days ago • 104
ASI-Bench: At the Dawn of Artificial Superintelligence Paper • 2608.17271 • Published 6 days ago • 60
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution Paper • 2608.16157 • Published 7 days ago • 80
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve Paper • 2608.16884 • Published 7 days ago • 17
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning Paper • 2607.29211 • Published 24 days ago • 14
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity Paper • 2608.13430 • Published 11 days ago • 13
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published 11 days ago • 54