AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling Paper • 2608.26623 • Published 11 days ago • 20
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments Paper • 2608.24804 • Published 13 days ago • 40
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published Jul 11 • 72
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory Paper • 2607.01071 • Published Jul 1 • 31
view article Article Hugging Face and Cerebras bring Gemma 4 to real-time voice AI +2 A-Mahla, andito, lvwerra, vyassaurabh • Jul 1 • 101
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics Paper • 2605.12178 • Published May 12 • 66
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents Paper • 2603.24440 • Published Mar 25 • 99
view article Article A New Framework for Evaluating Voice Agents (EVA) ServiceNow-AI • Mar 24 • 97
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings Paper • 2603.13594 • Published Mar 13 • 150
view article Article AprielGuard: A Guardrail for Safety and Adversarial Robustness in Modern LLM Systems ServiceNow-AI • Dec 23, 2025 • 50
ServiceNow-AI/Apriel-1.6-15b-Thinker Image-Text-to-Text • 15B • Updated Dec 22, 2025 • 29k • 305
view article Article Apriel-1.6-15b-Thinker: Cost-efficient Frontier Multimodal Performance ServiceNow-AI • Dec 9, 2025 • 84
Grounding Computer Use Agents on Human Demonstrations Paper • 2511.07332 • Published Nov 10, 2025 • 107
Surfer 2: The Next Generation of Cross-Platform Computer Use Agents Paper • 2510.19949 • Published Oct 22, 2025 • 38