BDH vs Linear-Attention Baselines β€” Replication Scaling Study

Independent replication of the scaling experiments from "Burst Denoising Hebbian Neural Networks" (Pathway, arXiv:2509.26507) and a head-to-head comparison of BDH-GPU against GPT-XL, GLA, DeltaNet, Mamba-2 at matched params, same tokens, two GPU replicas (RTX 4080 SUPER, A100-80GB).

Bottom line: BDH reaches consistently lower validation loss at 25M/50M/100M (~0.5–0.8 nats below the best baseline, stable across both GPUs).

Results β€” best validation loss (4080 / A100)

model 25M 50M 100M
BDH 2.5975/2.7917 2.6090/2.7555 2.7361/2.8117
GPT-XL 3.2441/3.3866 3.2476/3.4039 3.2842/3.3387
GLA 3.4040/3.5511 3.3767/3.5443 3.4605/3.6203
DeltaNet 3.3881/3.5556 3.3726/3.5684 3.4601/3.5705
Mamba-2 3.2357/3.4403 3.2351/3.3686 3.2360/3.3898

loss curves scaling

Protocol

  • Data: Europarl en-pl + en-cs (~1.28M pairs), byte-level UTF-8, <F:src>SRC<T:tgt>TGT format (paper Appendix B.1).
  • Training: AdamW lr 1e-3β†’1e-4, warmup 1000, wd 0.1, seq_len 2048 (100M: 1024), 4000 steps, TBPTT.
  • Models matched at equal total params; baselines use published fla kernels.
  • Full writeup: see PAPER.md.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for SlayerLab/bdh-scaling