Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Paper • 2608.20061 • Published 24 days ago • 46
McGill-NLP/LLM2Vec-Meta-Llama-3-8B-Instruct-mntp-supervised Sentence Similarity • Updated Apr 30, 2024 • 98.4k • 53
SmolLM2 Collection State-of-the-art compact LLMs for on-device applications: 1.7B, 360M, 135M • 16 items • Updated May 5, 2025 • 314
Magpie-Align/Llama-3-8B-Magpie-Align-SFT-v0.3 Text Generation • 8B • Updated Jul 19, 2024 • 8.54k • • 6