Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Paper • 2606.04923 • Published Jun 3 • 42
The Big Benchmarks Collection Collection Gathering benchmark spaces on the hub (beyond the Open LLM Leaderboard) • 13 items • Updated Nov 18, 2024 • 268
view article Article I trained a Language Model to schedule events with GRPO! anakin87 • Apr 29, 2025 • 95
Running on CPU Upgrade Agents 14 LLM Beer Distribution Game Public 👁 14 Play an interactive beer distribution game with AI
view article Article Cosmopedia: how to create large-scale synthetic data for pre-training Large Language Models +1 loubnabnl, anton-l, davanstrien • Mar 20, 2024 • 115
view article Article Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth mlabonne • Jul 29, 2024 • 373