CoolFace
Datasetpublicgated

m8than/cortex-20b

cortex-20b A 20 B-token (100 B-character), cap-balanced English/Japanese pretraining mixture assembled for general-purpose language model training. Release tag v0.9 (mixture build balanced-pretraining-v8, finalized 2026-09-15). The dataset is partitioned into three splits (train, validation, test) and stores 22.4 M text records on disk in zstd-compressed Parquet. The source stream spans web text, encyclopedias, news, public-domain books, academic articles, patents, court… See the full description on the dataset page: https://huggingface.co/datasets/m8than/cortex-20b.

sourceHugging Faceodc-byupdated 7d agoView on Hugging Face
0likes76downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

m8than/cortex-20b · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.