CoolFace
Datasetpublic

zee-drytis/damr-zyda2-64k-tokenized

DAMR Zyda2 64K Tokenized This repository contains a deterministic tokenized representation of a 1,749,978,785,641-token subset of Zyda-2 for language-model pretraining. Format Files under data/ contain contiguous little-endian uint16 token IDs. Concatenate shards in numeric order to reproduce the original stream. The shared 64K BPE tokenizer is stored under tokenizer/tokenizer.json. Shard manifests provide byte offsets, token counts, and SHA-256 checksums. The… See the full description on the dataset page: https://huggingface.co/datasets/zee-drytis/damr-zyda2-64k-tokenized.

sourceHugging Faceodc-byupdated 2mo agoView on Hugging Face
1likes370downloads
Dataset Card

DAMR Zyda2 64K Tokenized

This repository contains a deterministic tokenized representation of a 1,749,978,785,641-token subset of Zyda-2 for language-model pretraining.

Format

Files under data/ contain contiguous little-endian uint16 token IDs. Concatenate shards in numeric order to reproduce the original stream. The shared 64K BPE tokenizer is stored under tokenizer/tokenizer.json. Shard manifests provide byte offsets, token counts, and SHA-256 checksums.

The intended production mixture uses this source at weight 0.858486, code at 0.118510, and math at 0.023004.

License and attribution

Zyda-2 is distributed under ODC-By 1.0. Use remains subject to the terms of its underlying sources. Users must preserve attribution to Zyphra and the original Zyda-2 sources.

This tokenized representation adds no ownership claim over the underlying content.