datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commonwealth-wiki-mix
commonwealth-wiki-mix
FineWeb-style Parquet shards (LLM-training friendly) created by merging multiple Wikipedia language datasets into a single dataset to reduce looping during training.
What’s inside
Format: nanochat-parquet-v1
Layout: shard_*.parquet + metadata.json
Text column: text
Parquet settings: zstd (level 3), row_group_size=1024, use_dictionary=False, write_statistics=False
Sources included
This mix was built from the already-exported wiki datasets… See the full description on the dataset page: https://huggingface.co/datasets/JayJayThrowThrow/commonwealth-wiki-mix.nsw_commonwealth_corpus
