WikipediaGutenberg-0.5B
A from-scratch GPT-2-style language model (nanoGPT), 530M params, trained for 1 epoch on a combination of English Wikipedia (wikimedia/wikipedia, 20231101.en) and Project Gutenberg (common-pile/project_gutenberg_filtered) — ~9.0B tokens, ~17 tokens/parameter (near the Chinchilla compute-optimal ratio).
This is a research model from a study of how from-scratch pretraining behaves across capacity / data / exposure. It is not instruction-tuned, not safety-aligned, and not intended for production use.
Model details
- Architecture: GPT-2 (learned positional embeddings, GeLU MLP, full multi-head attention). 16 layers, n_embd 1536, 24 heads (head_dim 64) — 530.30M params.
- Training data: Wikipedia (4.17B tok) + Project Gutenberg (4.82B tok) = 8.99B train tokens (1 epoch); 1.00B val tokens.
- Optimizer: AdamW, LR 1.2e-3 → 1.2e-4 cosine, warmup 5000, batch 256 (65,536 tokens/iter), bf16, dropout 0.1.
- Hardware: 1× NVIDIA GH200, ~16.7 h.
- Status: training in progress.
ckpt.ptis the latest checkpoint (mirrored here every ~30 min); it carriesiter_num+ optimizer state for clean resumption.
⚠️ Format: nanoGPT, not HF Transformers
This is a raw nanoGPT checkpoint (a torch state_dict + config), not loadable via AutoModelForCausalLM.from_pretrained. Load it with nanoGPT:
import torch
from model import GPTConfig, GPT
ckpt = torch.load("ckpt.pt", map_location="cuda", weights_only=False)
model = GPT(GPTConfig(**ckpt["model_args"]))
model.load_state_dict(ckpt["model"])
Intended use & limitations
- Intended: ML research / education on pretraining dynamics and small-model knowledge elicitation.
- Not intended: deployment, factual QA, instruction-following. At 530M it has limited and weakly-elicitable factual knowledge (see the project's elicitation evals).
Evaluation
The full elicitation scoreboard (latent completion-probe recall, multiple-choice, free-form QA, few-shot/CoT) vs the 253M baselines will be added here on training completion.
Data & license
Wikipedia (CC BY-SA 4.0) + Project Gutenberg (public domain). Per Wikipedia's share-alike, model weights are released under CC BY-SA 4.0.
Sources
- nanoGPT: karpathy/nanoGPT
- Chinchilla (compute-optimal scaling): Hoffmann et al., 2022
- Datasets: wikimedia/wikipedia · common-pile/project_gutenberg_filtered