Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 1 day ago • 12
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 5 days ago • 7
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms Paper • 2609.23658 • Published 4 days ago • 11
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 1 day ago • 12
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 1 day ago • 12
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Paper • 2605.19436 • Published May 19 • 14
view article Article Navigating the RLHF Landscape: From Policy Gradients to PPO, GAE, and DPO for LLM Alignment NormalUhr • Feb 11, 2025 • 132
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device Paper • 2602.20161 • Published Feb 23 • 23