Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention Paper • 2606.20945 • Published Jun 18 • 80
Fortytwo-Network/Strand-Rust-Coder-14B-v1 Text Generation • 15B • Updated Jan 5 • 307 • • 183
Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding Paper • 2506.16035 • Published Jun 19, 2025 • 89
ZClip: Adaptive Spike Mitigation for LLM Pre-Training Paper • 2504.02507 • Published Apr 3, 2025 • 90
ZClip: Adaptive Spike Mitigation for LLM Pre-Training Paper • 2504.02507 • Published Apr 3, 2025 • 90
ZClip: Adaptive Spike Mitigation for LLM Pre-Training Paper • 2504.02507 • Published Apr 3, 2025 • 90 • 2
A Refined Analysis of Massive Activations in LLMs Paper • 2503.22329 • Published Mar 28, 2025 • 14
A Refined Analysis of Massive Activations in LLMs Paper • 2503.22329 • Published Mar 28, 2025 • 14
A Refined Analysis of Massive Activations in LLMs Paper • 2503.22329 • Published Mar 28, 2025 • 14 • 3
Variance Control via Weight Rescaling in LLM Pre-training Paper • 2503.17500 • Published Mar 21, 2025 • 5
Variance Control via Weight Rescaling in LLM Pre-training Paper • 2503.17500 • Published Mar 21, 2025 • 5
Variance Control via Weight Rescaling in LLM Pre-training Paper • 2503.17500 • Published Mar 21, 2025 • 5 • 2
Running 3.96k The Ultra-Scale Playbook 🌌 3.96k The ultimate guide to training LLM on large GPU Clusters
view article Article Fine-tuning LLMs with Singular Value Decomposition fractalego • Jun 2, 2024 • 14