TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 3 days ago • 144
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 6 days ago • 161
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation Paper • 2511.19365 • Published Nov 24, 2025 • 66
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation Paper • 2509.26391 • Published Sep 30, 2025 • 22
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation Paper • 2509.26391 • Published Sep 30, 2025 • 22
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation Paper • 2509.26391 • Published Sep 30, 2025 • 22 • 2
KS-Gen Collection Learning Human Skill Generators at Key-Step Levels • 3 items • Updated 7 days ago • 1
KS-Gen Collection Learning Human Skill Generators at Key-Step Levels • 3 items • Updated 7 days ago • 1
flateon/Belle-whisper-large-v3-turbo-zh-ct2 Automatic Speech Recognition • Updated Jan 17, 2025 • 41 • 4