bdh-scaling / README.md
kacperwikiel's picture
Upload folder using huggingface_hub
2f957bf verified
|
Raw
History Blame Contribute Delete
1.48 kB
---
language:
- en
- pl
- cs
license: apache-2.0
tags:
- bdh
- linear-attention
- scaling-laws
- replication
library_name: pytorch
---
# BDH vs Linear-Attention Baselines — Replication Scaling Study
Independent replication of the scaling experiments from **"Burst Denoising Hebbian Neural Networks"** (Pathway, arXiv:2509.26507) and a head-to-head comparison of **BDH-GPU** against **GPT-XL, GLA, DeltaNet, Mamba-2** at matched params, same tokens, two GPU replicas (RTX 4080 SUPER, A100-80GB).
**Bottom line:** BDH reaches consistently lower validation loss at 25M/50M/100M (~0.5–0.8 nats below the best baseline, stable across both GPUs).
## Results — best validation loss (4080 / A100)
| model | 25M | 50M | 100M |
|---|---|---|---|
| **BDH** | 2.5975/2.7917 | 2.6090/2.7555 | 2.7361/2.8117 |
| **GPT-XL** | 3.2441/3.3866 | 3.2476/3.4039 | 3.2842/3.3387 |
| **GLA** | 3.4040/3.5511 | 3.3767/3.5443 | 3.4605/3.6203 |
| **DeltaNet** | 3.3881/3.5556 | 3.3726/3.5684 | 3.4601/3.5705 |
| **Mamba-2** | 3.2357/3.4403 | 3.2351/3.3686 | 3.2360/3.3898 |
![loss curves](fig1_loss_curves_4080.png)
![scaling](fig2_scaling_both.png)
## Protocol
- Data: Europarl en-pl + en-cs (~1.28M pairs), byte-level UTF-8, `<F:src>SRC<T:tgt>TGT` format (paper Appendix B.1).
- Training: AdamW lr 1e-3→1e-4, warmup 1000, wd 0.1, seq_len 2048 (100M: 1024), 4000 steps, TBPTT.
- Models matched at equal total params; baselines use published fla kernels.
- Full writeup: see `PAPER.md`.