WikipediaGutenberg-0.5B

A from-scratch GPT-2-style language model (nanoGPT), 530M params, trained for 1 epoch on a combination of English Wikipedia (wikimedia/wikipedia, 20231101.en) and Project Gutenberg (common-pile/project_gutenberg_filtered) — ~9.0B tokens, ~17 tokens/parameter (near the Chinchilla compute-optimal ratio).

This is a research model from a study of how from-scratch pretraining behaves across capacity / data / exposure. It is not instruction-tuned, not safety-aligned, and not intended for production use.

Model details

  • Architecture: GPT-2 (learned positional embeddings, GeLU MLP, full multi-head attention). 16 layers, n_embd 1536, 24 heads (head_dim 64) — 530.30M params.
  • Training data: Wikipedia (4.17B tok) + Project Gutenberg (4.82B tok) = 8.99B train tokens (1 epoch); 1.00B val tokens.
  • Optimizer: AdamW, LR 1.2e-3 → 1.2e-4 cosine, warmup 5000, batch 256 (65,536 tokens/iter), bf16, dropout 0.1.
  • Hardware: 1× NVIDIA GH200, ~16.7 h.
  • Status: training in progress. ckpt.pt is the latest checkpoint (mirrored here every ~30 min); it carries iter_num + optimizer state for clean resumption.

⚠️ Format: nanoGPT, not HF Transformers

This is a raw nanoGPT checkpoint (a torch state_dict + config), not loadable via AutoModelForCausalLM.from_pretrained. Load it with nanoGPT:

import torch
from model import GPTConfig, GPT
ckpt = torch.load("ckpt.pt", map_location="cuda", weights_only=False)
model = GPT(GPTConfig(**ckpt["model_args"]))
model.load_state_dict(ckpt["model"])

Intended use & limitations

  • Intended: ML research / education on pretraining dynamics and small-model knowledge elicitation.
  • Not intended: deployment, factual QA, instruction-following. At 530M it has limited and weakly-elicitable factual knowledge (see the project's elicitation evals).

Evaluation

The full elicitation scoreboard (latent completion-probe recall, multiple-choice, free-form QA, few-shot/CoT) vs the 253M baselines will be added here on training completion.

Data & license

Wikipedia (CC BY-SA 4.0) + Project Gutenberg (public domain). Per Wikipedia's share-alike, model weights are released under CC BY-SA 4.0.

Sources

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for capotej/WikipediaGutenberg-0.5B