nanoGPT
Datasets
All datasets matching “nanoGPT”modded-nanogpt-kaon-track3-artifacts
Tokenbender Kaon Track 3 Artifacts
Artifact export from the 4x NVIDIA RTX PRO 6000 Blackwell pod cosmic-comet-e3 / tokenbender-kaon-track3.
This repository preserves run scripts, command lines, environment snapshots, GPU snapshots, status files, and logs. Generated torchinductor_cache directories and data/venv files are intentionally excluded.
Source GitHub issue: https://github.com/tokenbender/modded-nanogpt/issues/1
Source repo commit on pod:… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/modded-nanogpt-kaon-track3-artifacts.modded-nanogpt-noisenanogpt-tr-v5-data
nanogpt-tr-v5 Data
V5 (200M Türkçe LM) eğitimi için tokenize edilmiş veri.
Dosyalar
v5_stage1.bin — Web tier (OSCAR, mC4, forum, FineWeb-HQ) ~2.94B token
v5_stage2.bin — Medium tier (BellaTurca, Cosmos, CulturaX, Havadis, Cosmopedia) ~9.03B token
v5_stage3.bin — Premium tier (Wiki, Wikisource, Tezler, Akademik, FinePDFs, Özenli) ~2.97B token
v5_val.bin — Validation (3 stage'in son %1'i, ~150M token)
tokenizer-tr-v5.json — BPE tokenizer, 32K vocab, Stage3 üzerinde… See the full description on the dataset page: https://huggingface.co/datasets/musabc/nanogpt-tr-v5-data.tinystories-gpt2-nanogpt-bin
TinyStories GPT-2 Tokenized nanoGPT Shards
This dataset contains roneneldan/TinyStories tokenized with the GPT-2 tokenizer
and stored in the nanoGPT .bin format used by this repository.
Each .bin file contains:
a 256 int32 header
header[0] = 20240520
header[1] = 1
header[2] = number of uint16 tokens
GPT-2 token ids as uint16 values after the header
Files:
tinystories_train_000000.bin through tinystories_train_000004.bin
tinystories_val_000000.bin
.done marker files containing… See the full description on the dataset page: https://huggingface.co/datasets/hanspeterlyngsoeraaschoujensen/tinystories-gpt2-nanogpt-bin.openweb-nanogptnanogpt-tr-data
