SceneActBench: Can Agents Act on the 3D Scenes They See? Paper • 2607.22393 • Published 5 days ago • 6
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 5 days ago • 35
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 8 days ago • 72
FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry Paper • 2607.18227 • Published 9 days ago • 50
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 10 days ago • 165
From Pixels to States: Rethinking Interactive World Models as Game Engines Paper • 2607.14076 • Published 14 days ago • 35
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 15 days ago • 227
Phone Segmentation and Recognition through Phonological Activation Mapping Paper • 2607.09020 • Published 19 days ago • 9
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Paper • 2607.06559 • Published 22 days ago • 95
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space Paper • 2607.05373 • Published 23 days ago • 66
AnyBokeh: Physics-Guided Any-to-Any Bokeh Editing with Optical Fingerprint Transfer Paper • 2606.31959 • Published 28 days ago • 10
The Verification Horizon: No Silver Bullet for Coding Agent Rewards Paper • 2606.26300 • Published Jun 24 • 51