zee-drytis/damr-zyda2-64k-tokenized
DAMR Zyda2 64K Tokenized This repository contains a deterministic tokenized representation of a 1,749,978,785,641-token subset of Zyda-2 for language-model pretraining. Format Files under data/ contain contiguous little-endian uint16 token IDs. Concatenate shards in numeric order to reproduce the original stream. The shared 64K BPE tokenizer is stored under tokenizer/tokenizer.json. Shard manifests provide byte offsets, token counts, and SHA-256 checksums. The… See the full description on the dataset page: https://huggingface.co/datasets/zee-drytis/damr-zyda2-64k-tokenized.
DAMR Zyda2 64K Tokenized
This repository contains a deterministic tokenized representation of a 1,749,978,785,641-token subset of Zyda-2 for language-model pretraining.
Format
Files under data/ contain contiguous little-endian uint16 token IDs. Concatenate shards in numeric order to reproduce the original stream. The shared 64K BPE tokenizer is stored under tokenizer/tokenizer.json. Shard manifests provide byte offsets, token counts, and SHA-256 checksums.
The intended production mixture uses this source at weight 0.858486, code at 0.118510, and math at 0.023004.
License and attribution
Zyda-2 is distributed under ODC-By 1.0. Use remains subject to the terms of its underlying sources. Users must preserve attribution to Zyphra and the original Zyda-2 sources.
This tokenized representation adds no ownership claim over the underlying content.
