datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
latent_worker_early-a2_06latent_worker_early4_5latent_worker_early-a2_03latent_worker_early-a2_00latent_worker_early-a2_08latent_worker_early-a2_02latent_worker_early-a2_01latent_worker_early-a2_04latent_worker_early-a2_07latent_worker_early3_2latent_worker_early4_1latent_worker_early4_0latent_worker_early-a2_09latent_worker_early4_3latent_worker_early3_4latent_worker_early4_2latent_worker_early3_1latent_worker_early4_6latent_worker_early3_6latent_worker_early-a2_05latent_worker_early3_0small-llm-corpus-100b-v2-workers
Small-LM 100B — filtered English prose for small-model pretraining
100 billion tokens of English prose, filtered from nvidia/Nemotron-ClimbMix for a single purpose:
pretraining a small model where every token has to earn its place. No code, no LaTeX, no non-Latin
scripts, no menus or field lists — just prose, with a quality score attached to every document so
you can filter further without rebuilding.
Built by Edoardo and Rocco, equal authors.
Provenance and licence… See the full description on the dataset page: https://huggingface.co/datasets/roccoangelella/small-llm-corpus-100b-v2-workers.latent_worker_early4_7latent_worker_early-a1_02latent_worker_early-a1_00latent_worker_early-a1_09latent_worker_beta5_0latent_worker_early-a1_07latent_worker_early-a1_05latent_worker_early3_5
