Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI Paper • 2608.16319 • Published 7 days ago • 28
Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI Paper • 2608.16319 • Published 7 days ago • 28
Running 22 GLEE Competition — Live Leaderboard 🏆 22 Live leaderboard of the GLEE Competition @ NeurIPS 2026
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs Paper • 2604.14137 • Published Apr 16 • 10
Beyond IID: How General Are Tabular Foundation Models, Really? Paper • 2606.30410 • Published Jun 29 • 44
TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models Paper • 2511.08667 • Published Nov 11, 2025 • 7
LLM Explainability with Counterfactual Chains and Causal Graphs Paper • 2606.05972 • Published Jun 4 • 18
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks Paper • 2605.28556 • Published May 27 • 75
STRABLE: Benchmarking Tabular Machine Learning with Strings Paper • 2605.12292 • Published May 12 • 4
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image Paper • 2605.10616 • Published May 11 • 142
STRABLE: Benchmarking Tabular Machine Learning with Strings Paper • 2605.12292 • Published May 12 • 4
Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling Paper • 2605.12411 • Published May 12 • 49
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image Paper • 2605.10616 • Published May 11 • 142
Alignment Makes Language Models Normative, Not Descriptive Paper • 2603.17218 • Published Mar 17 • 46
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality Paper • 2602.14080 • Published Feb 15 • 23
STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts Paper • 2602.14265 • Published Feb 15 • 21