CoolFace
Datasetpublicgated

m8than/cortex-20b

cortex-20b A 20 B-token (100 B-character), cap-balanced English/Japanese pretraining mixture assembled for general-purpose language model training. Release tag v0.9 (mixture build balanced-pretraining-v8, finalized 2026-09-15). The dataset is partitioned into three splits (train, validation, test) and stores 22.4 M text records on disk in zstd-compressed Parquet. The source stream spans web text, encyclopedias, news, public-domain books, academic articles, patents, court… See the full description on the dataset page: https://huggingface.co/datasets/m8than/cortex-20b.

sourceHugging Faceodc-byupdated 7d agoView on Hugging Face
0likes76downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.