Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 6 days ago • 298
Laguna XS 2.1 Collection Designed for agentic coding and long-horizon work on a local machine. Licensed under OpenMDW-1.1. • 9 items • Updated Jul 2 • 23
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published 9 days ago • 35
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 13 days ago • 39
Distilled Reinforcement Learning for LLM Post-training Paper • 2607.17247 • Published 17 days ago • 9
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 20 days ago • 171
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation Paper • 2607.13124 • Published 22 days ago • 20
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Paper • 2607.12395 • Published 22 days ago • 99
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 18 days ago • 138
view article Article Native-speed vLLM transformers modeling backend hmellor, lysandre • 28 days ago • 65
Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model Paper • 2607.03509 • Published Jul 3 • 14
RADIO Collection A collection of Foundation Vision Models that combine multiple models (CLIP, DINOv2, SAM, etc.). • 19 items • Updated 19 days ago • 38
Wan-Streamer v0.2: Higher Resolution, Same Latency Paper • 2607.04443 • Published about 1 month ago • 40
DataComp-VLM: Improved Open Datasets for Vision-Language Models Paper • 2606.28551 • Published Jun 26 • 52
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing Paper • 2606.26740 • Published Jun 25 • 82
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Paper • 2606.27313 • Published Jun 25 • 38