OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published 11 days ago • 67
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 4 days ago • 42
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 5 days ago • 20
MiniWorld: Democratizing the Training of Video World Models from Scratch Paper • 2608.01127 • Published 8 days ago • 19
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published 10 days ago • 23