Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Paper • 2607.01191 • Published 22 days ago • 19
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement Paper • 2607.00446 • Published 22 days ago • 23
WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory Paper • 2607.02517 • Published 21 days ago • 33
MemLearner: Learning to Query Context memory for Video World Models Paper • 2606.31734 • Published 23 days ago • 28
PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled Denoising Paper • 2606.30968 • Published 24 days ago • 28
Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling Paper • 2607.01642 • Published 21 days ago • 39
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Paper • 2607.02255 • Published 21 days ago • 64
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories Paper • 2606.11176 • Published Jun 9 • 132
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning Paper • 2602.10090 • Published Feb 10 • 53
Running on Zero Agents Featured 174 SoulX-Singer 🎤 174 Generate singing voice from lyrics and convert vocals
Running Agents Featured 59 ERNIE-4.5-VL-28B-A3B-Thinking Demo 👐 59 Compact model, powerful multimodal reasoning.