SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 6 days ago • 93
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Paper • 2607.07740 • Published 14 days ago • 23
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 13 days ago • 74
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published 9 days ago • 16
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published 19 days ago • 79
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory Paper • 2607.01071 • Published 21 days ago • 30
Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention Paper • 2606.20945 • Published Jun 18 • 80
Learning from the Self-future: On-policy Self-distillation for dLLMs Paper • 2606.18195 • Published Jun 16 • 77