view article Article BenchMIRT: What are LLM benchmarks actually measuring? allenai • 18 days ago • 23
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 17 days ago • 125
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 16 days ago • 107
view article Article Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI nico-martin, Xenova • 19 days ago • 72
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • Aug 14 • 206