DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 6 days ago • 171 • 4
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 27 days ago • 40 • 7
SimpleGPT: Improving GPT via A Simple Normalization Strategy Paper • 2602.01212 • Published Feb 1 • 4 • 6
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels Paper • 2602.11715 • Published Feb 12 • 8 • 3
Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm Paper • 2602.11543 • Published Feb 12 • 6 • 4
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models Paper • 2602.06694 • Published Feb 6 • 22 • 8
SimpleGPT: Improving GPT via A Simple Normalization Strategy Paper • 2602.01212 • Published Feb 1 • 4 • 6
OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer Paper • 2601.14250 • Published Jan 20 • 48 • 5