BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender Paper β’ 2609.15478 β’ Published 5 days ago β’ 24
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination Paper β’ 2511.17490 β’ Published Nov 21, 2025 β’ 22
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models Paper β’ 2510.05034 β’ Published Oct 6, 2025 β’ 51
view article Article Illustrating Reinforcement Learning from Human Feedback (RLHF) +2 natolambert, LouisCastricato, lvwerra, Dahoas β’ Dec 9, 2022 β’ 426
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness Paper β’ 2505.20426 β’ Published May 26, 2025 β’ 7
Running on Zero Agents Featured 2.05k Stable Diffusion 3.5 Large π 2.05k Generate images with SD3.5
IDEA-Research/grounding-dino-tiny Zero-Shot Object Detection β’ 0.2B β’ Updated May 12, 2024 β’ 763k β’ 116
meta-llama/Llama-3.2-11B-Vision Image-Text-to-Text β’ 11B β’ Updated Sep 27, 2024 β’ 11.6k β’ 606
Running on Zero Agents Featured 614 Unofficial SDXL Turbo Img2Img Txt2Img π¬ 614 Generate images from text or photos in realβtime