m8than/cortex-20b
cortex-20b A 20 B-token (100 B-character), cap-balanced English/Japanese pretraining mixture assembled for general-purpose language model training. Release tag v0.9 (mixture build balanced-pretraining-v8, finalized 2026-09-15). The dataset is partitioned into three splits (train, validation, test) and stores 22.4 M text records on disk in zstd-compressed Parquet. The source stream spans web text, encyclopedias, news, public-domain books, academic articles, patents, court… See the full description on the dataset page: https://huggingface.co/datasets/m8than/cortex-20b.
Nothing at this path on main. The folder may be empty, or the revision may not exist.
This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.
