TimeLens Collection [CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs • 5 items • Updated Feb 24 • 11
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 9 days ago • 165
VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance Paper • 2607.14660 • Published 12 days ago • 9
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 12 days ago • 169
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation Paper • 2606.26016 • Published Jun 24 • 10
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer Paper • 2606.16255 • Published Jun 15 • 15
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Paper • 2606.13289 • Published Jun 11 • 30
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision Paper • 2512.01342 • Published Dec 1, 2025 • 21
SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation Paper • 2511.19320 • Published Nov 24, 2025 • 43
Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment Paper • 2511.22345 • Published Nov 27, 2025 • 13
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training Paper • 2203.12602 • Published Mar 23, 2022 • 6
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models Paper • 2504.15271 • Published Apr 21, 2025 • 69
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions Paper • 2511.03334 • Published Nov 5, 2025 • 54