AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published 11 days ago • 44
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence Paper • 2607.06838 • Published 19 days ago • 14
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published 18 days ago • 64
Scalable Visual Pretraining for Language Intelligence Paper • 2607.09657 • Published 16 days ago • 56
RNGBench Collection Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games • 2 items • Updated Jun 23 • 3
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Paper • 2606.19338 • Published Jun 17 • 50
N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization Paper • 2606.10768 • Published Jun 9 • 25
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Paper • 2605.13831 • Published May 13 • 90
OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Paper • 2604.18486 • Published Apr 20 • 96
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Paper • 2604.18292 • Published Apr 20 • 89
Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization Paper • 2603.28342 • Published Mar 30 • 26
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale Paper • 2603.25040 • Published Mar 26 • 134
DataChef: Cooking Up Optimal Data Recipes for LLM Adaptation via Reinforcement Learning Paper • 2602.11089 • Published Feb 11 • 18