view article Article LFM2.5-Encoders for Fast Long-Context Inference on CPU LiquidAI • 7 days ago • 65
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale Paper • 2505.03005 • Published May 5, 2025 • 36
view article Article Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 18 days ago • 191
view article Article Bringing Nunchaku 4-bit Diffusion Inference to Diffusers rootonchair, sayakpaul • 13 days ago • 63
Instella: Fully Open Language Models with Stellar Performance Paper • 2511.10628 • Published Nov 13, 2025 • 6
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing Paper • 2607.08646 • Published 27 days ago • 3
KronQ: LLM Quantization via Kronecker-Factored Hessian Paper • 2607.07964 • Published 28 days ago • 32
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 26 days ago • 86
Multiplayer Interactive World Models with Representation Autoencoders Paper • 2607.05352 • Published 30 days ago • 28
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published 28 days ago • 64
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published Jul 3 • 84
Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation Paper • 2607.06957 • Published 28 days ago • 12