Abstract
Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them with a core-cognition inspired dataset? We introduce WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories. We build Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure, yielding 10,000+ samples per task. We release a 1.5M-sample training corpus and a 300-question exam. On this exam we evaluate 14 video models: 3 reference-to-video, 7 edit, and 4 continuation, among which PWM-WROP, our 16B world model. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models. We release the data, exam, model answers, scores, weights, and PWM, our native-PyTorch training stack on AWS Trainium2.
Community
Object permanence is the foundation of human cognition. Here we present a very complete data infrastructure that's composed of a very diverse set of object permanence cognitive tasks, and with each task we have a Blender-based data generator that allows one to scale each task to at least 10,000 diverse data samples. We have shown the effectiveness of this data infrastructure in training video models and world models with object permanence.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- VGI-Bench: Probing Visual Intelligence in Video Generation Models (2026)
- Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation (2026)
- Can 4D Foundation Models Remember? (2026)
- Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEval (2026)
- WorldReward: Reward Modeling for Camera-Conditioned World Models (2026)
- From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models (2026)
- DynaPix: Can Vision-Language Models Identify the Exact Future? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.28654 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 2
Hokin/object-permanence
Hokin/object-permanence-benchmark
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper