CoolFace
Modelpublic

ctxnn1/gpt2-124m-fineweb-edu-10b

sourceHugging Facemitupdated 2mo agoView on Hugging Face
1likes53downloads
Model Card

GPT-2 124M FineWeb-Edu 10B

This is a GPT-2 124M base text-completion model trained from scratch on the FineWeb-Edu sample-10BT dataset. It is not instruction-tuned and is not a chatbot.

Try it live

Run this model in a free public Gradio Space:

[Open the live demo →](https://huggingface.co/spaces/ctxnn1/gpt2-from-scratch-demo)

The Space runs on CPU-basic (free) hardware and accepts a prompt with temperature, top-k, top-p, repetition penalty, and seed controls.

Model description

  • —Architecture: GPT-2 decoder-only Transformer
  • —Parameters: 124,439,808 after export-vocabulary trimming
  • —Layers / heads / hidden size: 12 / 12 / 768
  • —Context length: 1,024 tokens
  • —Export vocabulary: 50,257 GPT-2 tokens
  • —Training tokens: 9,999,745,024
  • —Hardware: one NVIDIA H100

The native trainer padded its embedding/output matrix from 50,257 to 50,304 rows for efficient kernels. The final 47 rows were never tokenizer-addressable. This export keeps rows 0–50,256, sets config.vocab_size=50257, and preserves tied input/output embeddings, so generation cannot emit a padded ID.

Final evaluation

MetricValue
Training step19,073
Train loss3.103327
Validation loss3.030832
Validation perplexity20.714451
HellaSwag accuracy30.0339% (3,016/10,042)

Intended use

The model is intended for research, education, reproducibility studies, and experiments with small pretrained language models. It performs ordinary next-token completion.

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "ctxnn1/gpt2-124m-fineweb-edu-10b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
inputs = tokenizer("The future of artificial intelligence is", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=64, do_sample=True, top_k=50)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Limitations, risks, and biases

  • —This is a 124M-parameter base model and is not competitive with modern large language models.
  • —It is not instruction-tuned, preference-aligned, safety-tuned, or suitable as a chatbot.
  • —Outputs can be inaccurate, incoherent, biased, offensive, unsafe, or memorized from pretraining data.
  • —HellaSwag accuracy is only modestly above the 25% random-choice baseline.
  • —The run did not include comprehensive safety, fairness, memorization, or downstream-task evaluation.
  • —Users must evaluate outputs and suitability for their own domain before deployment.

Training and conversion provenance

  • —Source code: https://github.com/ctxnn/gpt-2
  • —W&B: https://wandb.ai/ctxnn-thapar-university/gpt2-from-scratch/runs/65e78f54c14046ef99e04e12e7b3e810
  • —Native checkpoint SHA-256: e519d993d20c98c841ef061f76a1dec3e6ee24d5e55162bdea2a3e2da280fd40
  • —Training Git SHA: 36bfc9edd044eb828e118d49c79532eef8440a2a
  • —Cloud execution Git SHA: 9862792ab4024f9ebec758be73ebe7e75419d09b

The native Linear weights for attention and MLP projections were transposed into Hugging Face GPT-2 Conv1D orientation. Positional embeddings, LayerNorm parameters, attention/MLP projections, and tied token embeddings were preserved and validated with native-versus-Hugging-Face logit and loss comparisons.