zee-drytis/damr-zyda2-64k-tokenized
DAMR Zyda2 64K Tokenized This repository contains a deterministic tokenized representation of a 1,749,978,785,641-token subset of Zyda-2 for language-model pretraining. Format Files under data/ contain contiguous little-endian uint16 token IDs. Concatenate shards in numeric order to reproduce the original stream. The shared 64K BPE tokenizer is stored under tokenizer/tokenizer.json. Shard manifests provide byte offsets, token counts, and SHA-256 checksums. The… See the full description on the dataset page: https://huggingface.co/datasets/zee-drytis/damr-zyda2-64k-tokenized.
This repository belongs to zee-drytis on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
