Text Generation
Transformers
Safetensors
code
gpt2
text-generation-inference

TULLUS/codeparrot-small-multi 🦜

A clean, verified Safetensors conversion of codeparrot/codeparrot-small-multi.

This repository preserves the original model weights while normalizing the legacy GPT-2 embedding representation from two physically duplicated matrices into properly tied input/output embeddings.

What this is

This is a checkpoint conversion, not a new training run.

The original CodeParrot checkpoint was loaded from its legacy pytorch_model.bin, inspected, normalized, saved as Safetensors, reloaded, and then subjected to exact tensor verification.

The goal was to produce a clean modern Transformers checkpoint without changing the learned weight values.

Source

Original model: codeparrot/codeparrot-small-multi

Original architecture: GPT-2 / GPT2LMHeadModel

Original tokenizer: GPT-2 tokenizer

Vocabulary size: 32,768

BOS token ID: 0

EOS token ID: 0

PAD token: none

Original checkpoint SHA256:

0207f6b427e3cbf1bcb9726abb6bbba6620e9912e1abf23087ccdef61818ffb2

Conversion details

The legacy checkpoint contained two separate physical copies of the embedding matrix:

transformer.wte.weight
lm_head.weight

Both tensors were verified to have identical values:

shape:            (32768, 768)
exact equality:   True
max difference:   0.0
same storage:     False

The model was then normalized with:

tie_word_embeddings = True

After normalization:

same storage:      True
exact equality:    True

The resulting checkpoint therefore uses a single shared embedding tensor for the input embeddings and language-model head.

Parameter count

The legacy checkpoint physically stored:

136,174,080 parameters

This included the duplicated embedding storage.

After tying the identical embeddings, the clean model contains:

111,008,256 unique parameters

The resulting Safetensors file is approximately:

444,048,000 bytes
423.48 MiB

Verification

The conversion was validated after saving and reloading the Safetensors checkpoint.

Verified checks include:

  • Original source SHA256
  • Source tokenizer length
  • BOS/EOS token IDs
  • Source embedding values
  • Source physical parameter count
  • Embedding storage normalization
  • Safetensors serialization
  • Reload into GPT2LMHeadModel
  • tie_word_embeddings=True
  • Final tied embedding storage
  • Metadata consistency
  • Exact tensor equality

Final verification result:

Weights: EXACTLY PRESERVED
Embedding configuration: TIED
Unique parameter count: 111,008,256
BOS/EOS IDs: 0 / 0
Exact tensor verification: PASS

The final comparison verified all 149 logical model tensors after reload. Safetensors physically stores 148 tensors because the tied lm_head.weight is represented by the shared transformer.wte.weight storage.

About the conversion warning

During loading of the original legacy checkpoint, Transformers reported the following unexpected entries:

transformer.h.{0...11}.attn.bias
transformer.h.{0...11}.attn.masked_bias

These are legacy GPT-2 attention-mask buffers rather than learned model weights. They do not represent missing trained parameters, and the complete learned tensor set was subsequently verified exactly after conversion and reload.

Intended use

This checkpoint is intended as a clean starting point for experimentation with the CodeParrot GPT-2 architecture, including further fine-tuning and research experiments.

It is also the clean base checkpoint for the planned:

TULLUS-CodeParrot-???

training experiment.

Example usage

from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "TULLUS/codeparrot-small-multi"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

prompt = "def fibonacci(n):"

inputs = tokenizer(prompt, return_tensors="pt")

outputs = model.generate(
    **inputs,
    max_new_tokens=128,
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Provenance

This repository is a conversion of the original:

codeparrot/codeparrot-small-multi

No claim is made here that the converted checkpoint is a new trained model. The purpose of this repository is to provide a clean Safetensors representation with normalized tied embeddings while preserving the original learned values.


TULLUS BAKES 🍪

Fresh From The AI Ovens, It's The Goods..

Downloads last month
104
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TULLUS/codeparrot-small-multi

Finetuned
(1)
this model

Datasets used to train TULLUS/codeparrot-small-multi

Collection including TULLUS/codeparrot-small-multi