τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction Paper • 2609.04611 • Published 5 days ago • 5
τ-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge Paper • 2603.04370 • Published Mar 4 • 3
τ-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge Paper • 2603.04370 • Published Mar 4 • 3
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration Paper • 2506.05579 • Published Jun 5, 2025 • 4
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration Paper • 2506.05579 • Published Jun 5, 2025 • 4 • 2
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval Paper • 2407.12883 • Published Jul 16, 2024 • 13
IMPersona: Evaluating Individual Level LM Impersonation Paper • 2504.04332 • Published Apr 6, 2025 • 2
IMPersona: Evaluating Individual Level LM Impersonation Paper • 2504.04332 • Published Apr 6, 2025 • 2
view article Article Introducing the LiveCodeBench Leaderboard - Holistic and Contamination-Free Evaluation of Code LLMs +5 StringChaos, minimario, tianjunz, xu3kev, kingh0730, FanjiaYan, clefourrier • Apr 16, 2024 • 16