SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 3 days ago • 56
Environment-free Synthetic Data Generation for API-Calling Agents Paper • 2607.16900 • Published 7 days ago • 20
SciForma: Structure-Faithful Generation of Scientific Diagrams Paper • 2607.18091 • Published 5 days ago • 22
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 9 days ago • 168
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation Paper • 2607.13124 • Published 11 days ago • 19
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 9 days ago • 99
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning Paper • 2607.08393 • Published 15 days ago • 18
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models Paper • 2607.08770 • Published 16 days ago • 37
view article Article Profiling in PyTorch (Part 3): Attention is all you profile +2 ariG23498, sergiopaniego, sayakpaul, ror • 15 days ago • 38
Unified Audio Intelligence Without Regressing on Text Intelligence Paper • 2607.05196 • Published 19 days ago • 23
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning Paper • 2607.04425 • Published 20 days ago • 72
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published 22 days ago • 81
view article Article Atom2.7m: Representation-Level Specialization for Arithmetic-Aware Small Language Models ucr-max • 17 days ago • 9
Kronos: A Foundation Model for the Language of Financial Markets Paper • 2508.02739 • Published Aug 2, 2025 • 48