MBZUAI/MediX-R1-2B-GGUF
Image-Text-to-Text • 2B • Updated • 486 • 1
Natural Language Processing, Machine Learning, and Computer Vision
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Training-Free Speech-Centric Omni Understanding with Frozen VLMs