CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceBio /carbon-pretraining-corpus 🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.tabulartext-generation100M<n<1B30 likes3.6k downloads3mo agoHugging Face02Kiy-K /pretraining-corpus 🧠 Kiy-K Synthetic Pretraining Corpus Author: Khoi K. (@Kiy-K)License: Apache 2.0Last Updated: 2025-10-30 📘 Overview The Kiy-K Synthetic Pretraining Corpus is a large-scale collection of synthetically generated English text designed for language model pretraining and instruction-tuning research. All data is synthetic, created using open-source large language models such as GPT-OSS, NVIDIA Nemotron, and DeepSeek, under full control of the author.No real user… See the full description on the dataset page: https://huggingface.co/datasets/Kiy-K/pretraining-corpus.tabulartext-generation10K<n<100K3 likes386 downloads10mo agoHugging Face03JustACluelessKidAtSchool /tiny-slm-pretraining-corpus 🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB) A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures. 100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders. 📊 Dataset Statistics Total Documents: 20,066,075 Train: 19,663,898 Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tabulartext-generation10M<n<100M0 likes198 downloads1mo agoHugging Face04lapa-llm /pretraining-high-quality Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness of the text lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.tabulartext-generation10M<n<100M0 likes139 downloads10mo agoHugging Face05lapa-llm /pretraining-lower-quality Dataset Card for Lapa Pretraining Lower Quality Dataset Dataset Description Dataset Summary This dataset is a high quality (but lower quality than https://huggingface.co/datasets/lapa-llm/pretraining-high-qualit) subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality.tabulartext-generation10M<n<100M0 likes102 downloads10mo agoHugging Face06gugarosa /synthetic-pretraining-transformers-v1 Synthetic Pre-training Transformers v1.0.0 Dataset Description This is a synthetic pre-training dataset generated from transformer architecture patterns. It contains paraphrased, augmented, and interpolated content derived from validated seed data about neural sequence modeling and attention mechanisms. Dataset Summary Total Samples: 100 Total Tokens: 6,084 Average Tokens per Sample: 60.84 Format: Parquet Version: 1.0.0 License: CC-BY-4.0 Supported… See the full description on the dataset page: https://huggingface.co/datasets/gugarosa/synthetic-pretraining-transformers-v1.tabulartext-generationn<1K0 likes26 downloads6mo agoHugging Face07blab-jhu /KYS-1.5B-Pretraining-Corporagated KYS-1.5B-Pretraining-Corpora The six 10B-token pretraining mixtures from Know Your Sources: Data Selection Matters when Rewriting for Data-Constrained Pretraining, stored without duplication. The key idea: one anchor, six remainders Every setting trains on the same 10B-token recipe: 10B mixture = 5B shared anchor + 5B strategy-specific tokens (identical in all (this is the ONLY thing six settings, that differs… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-1.5B-Pretraining-Corpora.tabulartext-generation10M<n<100M0 likes19 downloads1mo agoHugging Face08transhumanist-already-exists /pretraining-high-quality-10k-workshop Lapa HQ 10k Workshop Corpus A small deterministic subset of lapa-llm/pretraining-high-quality for tokenizer-transfer workshop runs. Provenance Source dataset: lapa-llm/pretraining-high-quality Source config: default Source split: train Rows: 10000 Selection: first 10000 rows by dataset-server row order Download window size: 100 Parallel workers: 20 Created at UTC: 2026-06-20T09:22:21.460168+00:00 Added columns: source_row_idx mini_corpus_index tabulartext-generation10K<n<100K0 likes17 downloads3mo agoHugging Face09fffoivos /glossapi-greek-nanochat-pretraining-datasetgated Glossapi Greek Nanochat Pretraining Dataset This repository contains the source-separated Greek corpus used to build nanochat Greek pretraining mixtures. It is intentionally not a pre-split train/validation/test export: builders load data/*.parquet, preserve source_dataset, and create deterministic experiment-specific mixes and splits downstream. Current Snapshot Total rows: 49474947 Total characters: 248276390721 Included source datasets: 19 Data files: 273… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset.tabulartext-generation10M<n<100M0 likes5 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.