SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 7 days ago • 95
UniVR: Thinking in Visual Space for Unified Visual Reasoning Paper • 2607.12800 • Published 9 days ago • 32
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published 25 days ago • 168
Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization Paper • 2605.28109 • Published May 27 • 23
AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents Paper • 2603.27490 • Published Mar 29 • 20
KV-Embedding: Training-free Text Embedding via Internal KV Re-routing in Decoder-only LLMs Paper • 2601.01046 • Published Jan 3 • 14