Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss Paper • 2608.03796 • Published 7 days ago • 9 • 2
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces Paper • 2608.03451 • Published 7 days ago • 30 • 4
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay Paper • 2608.05784 • Published 5 days ago • 24 • 6
Invisible Shortcuts: Why Vision Encoders Know Your Camera Paper • 2608.05424 • Published 6 days ago • 16 • 3
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents Paper • 2608.05013 • Published 7 days ago • 34 • 3
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? Paper • 2608.03874 • Published 7 days ago • 14 • 4
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants Paper • 2607.26611 • Published 13 days ago • 32 • 3
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space Paper • 2607.25675 • Published 14 days ago • 65 • 3
A New Role for Relevance: Guiding Corpus Interaction in Agentic Search Paper • 2607.24223 • Published 15 days ago • 94 • 6
Codifying the Judge: Scalable Evaluation via Program Distillation Paper • 2607.22561 • Published May 29 • 8 • 3
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems Paper • 2607.21503 • Published 19 days ago • 28 • 4
OpenForgeRL: Train Harness-native Agents in Any Environment Paper • 2607.21557 • Published 19 days ago • 9 • 4
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 20 days ago • 30 • 6
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents Paper • 2607.18754 • Published 21 days ago • 25 • 4
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment Paper • 2607.07820 • Published Jul 8 • 92 • 4
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM Paper • 2607.11683 • Published 29 days ago • 148 • 4
Rethinking the Evaluation of Harness Evolution for Agents Paper • 2607.12227 • Published 28 days ago • 10 • 3