Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? Paper • 2609.10226 • Published 2 days ago • 8
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 3 days ago • 386
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 8 days ago • 137
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published 7 days ago • 20
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp Image-Text-to-Text • 305B • Updated 10 days ago • 401k • • 858
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 8 days ago • 291
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 10 days ago • 100
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 11 days ago • 58
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation Paper • 2607.05147 • Published Jul 6 • 49
gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 Image-Text-to-Text • 15B • Updated about 20 hours ago • 714k • 153