SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL Paper • 2609.29050 • Published 7 days ago • 10
Block Sparse Attention with Log-Linear Complexity Paper • 2609.31093 • Published 6 days ago • 27
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures Paper • 2609.29429 • Published 7 days ago • 22
Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery Paper • 2609.27980 • Published 8 days ago • 6
Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models Paper • 2609.26637 • Published 9 days ago • 23
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 7 days ago • 99
GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression Paper • 2609.25963 • Published 9 days ago • 17
FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation Paper • 2609.27657 • Published 8 days ago • 9
Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone Paper • 2609.23087 • Published 12 days ago • 11
MemoryAthena: Adaptive Routing over Latent and Generated Memories Paper • 2609.25853 • Published 9 days ago • 11
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms Paper • 2609.27321 • Published 8 days ago • 22
KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling Paper • 2609.27294 • Published 8 days ago • 1
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 9 days ago • 36
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization Paper • 2609.15991 • Published 13 days ago • 17
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 12 days ago • 17
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Paper • 2609.25804 • Published 9 days ago • 160
Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention Paper • 2609.24797 • Published 10 days ago • 11
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 10 days ago • 55