CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01applied-ai-018 /pretraining_v1-omega_bookstabular100M<n<1B25 likes421k downloads2y agoHugging Face02HuggingFaceBio /carbon-pretraining-corpus 🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.tabulartext-generation100M<n<1B30 likes3.7k downloads3mo agoHugging Face03EleutherAI /filtering-pretraining-mix-arrow-formattabular100M<n<1B0 likes2.9k downloads2y agoHugging Face04JuaAI /ts-icl-pretraining-corpus TS-ICL Pretraining Corpus (community reconstruction) A unified, cleaned reconstruction of the univariate pretraining corpus described in Table 5 of TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning (Le Naour, Nabil & Petralia, EDF R&D; arXiv:2606.05878). The TS-ICL authors did not release their pretraining data pipeline, so this corpus is rebuilt from the named upstream sources (LOTSA, Chronos, and the TempoPFN synthetic generators) and… See the full description on the dataset page: https://huggingface.co/datasets/JuaAI/ts-icl-pretraining-corpus.tabulartime-series-forecasting1M<n<10M0 likes2k downloads3mo agoHugging Face05avewright /tabula-pretraining-corpus-v2 Tabula Pretraining Corpus v2 A large-scale synthetic tabular dataset for pretraining transformer-based in-context learning models for tabular data (similar to TabPFN). Overview Metric Value Total rows 272,271,776 Total datasets 10,867 Shards 135 Mean utility AUC 0.851 Format Parquet (float32) Schema Each shard is a Parquet file with a fixed-width schema: feat_0 through feat_63: Float32 feature columns. Unused slots are NaN. target:… See the full description on the dataset page: https://huggingface.co/datasets/avewright/tabula-pretraining-corpus-v2.tabulartabular-classification1B<n<10B0 likes1.9k downloads6mo agoHugging Face06aklein4 /seq2seq-mixed-pretraining-SmolLM2tabular100M<n<1B1 likes1.6k downloads8mo agoHugging Face07EleutherAI /deep-ignorance-pretraining-mix Deep Ignorance Model Suite We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-pretraining-mix.tabular100M<n<1B4 likes926 downloads1y agoHugging Face08nasa-ibm-ai4science /Sombench-pretraining-data SomBench Pre-training Corpus: Multimodal Lunar Tiles Dataset Summary This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining. Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.tabularn<1K0 likes469 downloads13d agoHugging Face09orionweller /contrastive-pretraining Contrastive Pretraining Per-language query/document pairs produced by the retrieval-common-crawl pipeline. Each config corresponds to a single language or source with identical LightOn-style schema. Config overview Configs available are: fw-edu, fw2-arb_Arab, fw2-ces_Latn, fw2-cmn_Hani, fw2-dan_Latn, fw2-deu_Latn, fw2-ell_Grek, fw2-fas_Arab, fw2-fra_Latn, fw2-hun_Latn, fw2-ind_Latn, fw2-ita_Latn, fw2-jpn_Jpan, fw2-nld_Latn, fw2-pol_Latn, fw2-por_Latn, fw2-rus_Cyrl… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/contrastive-pretraining.tabular100M<n<1B3 likes463 downloads5mo agoHugging Face10aklein4 /mixed-pretraining-tokenizedtabular10M<n<100M0 likes431 downloads1y agoHugging Face11neuralbioinfo /ncbi-dataset-for-genome-network-pretrainingtabular100M<n<1B0 likes418 downloads5mo agoHugging Face12sudoers /control-pretraining-filter-annotatedgated control-pretraining-filter-annotated climbmix_full — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content) climbmix_ai_docs — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by content) climbmix_long — from geodesic-research/control-pretraining-datasets (stage decisions reused from the run-1 dataset where documents match by… See the full description on the dataset page: https://huggingface.co/datasets/sudoers/control-pretraining-filter-annotated.tabular100M<n<1B0 likes410 downloads13h agoHugging Face13Kiy-K /pretraining-corpus 🧠 Kiy-K Synthetic Pretraining Corpus Author: Khoi K. (@Kiy-K)License: Apache 2.0Last Updated: 2025-10-30 📘 Overview The Kiy-K Synthetic Pretraining Corpus is a large-scale collection of synthetically generated English text designed for language model pretraining and instruction-tuning research. All data is synthetic, created using open-source large language models such as GPT-OSS, NVIDIA Nemotron, and DeepSeek, under full control of the author.No real user… See the full description on the dataset page: https://huggingface.co/datasets/Kiy-K/pretraining-corpus.tabulartext-generation10K<n<100K3 likes368 downloads10mo agoHugging Face14fffoivos /glossapi-greek-nanochat-pretraining-dataset-v2 GlossAPI Greek pretraining corpus v2 HPLT filtering method The HPLT component is HPLT/ell_Grek_ge8_no_mt_clean60. It retains HPLT quality bins 8, 9, 10 (GE8), uses the pre-applied no-MT/register filter, requires greek_badness_score <= 60, and applies Wave4 Greek re-cleaning and normalization. The standalone filtered slice contained 48,728,774 documents; 48,629,460 remain after corpus-wide deduplication. GlossAPI datasets and token counts GlossAPI… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset-v2.tabular10M<n<100M0 likes254 downloads1mo agoHugging Face15fineinstructions-pretraining /ipt_fineinstructions_all If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } tabular10M<n<100M0 likes246 downloads8mo agoHugging Face16JustACluelessKidAtSchool /tiny-slm-pretraining-corpus 🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB) A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures. 100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders. 📊 Dataset Statistics Total Documents: 20,066,075 Train: 19,663,898 Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tabulartext-generation10M<n<100M0 likes224 downloads1mo agoHugging Face17continuallearning /pretraining_v1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "franka", "total_episodes": 969, "total_frames": 259815, "total_tasks": 6, "total_videos": 1938, "total_chunks": 1, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:969" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/pretraining_v1.tabularrobotics100K<n<1M0 likes218 downloads6mo agoHugging Face18avewright /tabula-pretraining-corpus Tabula Pretraining Corpus A continuously growing tabular pretraining corpus for the Tabula foundation model (tabPFN-style in-context learning). Built by an autonomous agent that alternates between harvesting permissively-licensed real datasets and generating high-quality synthetic ones. Usage from datasets import load_dataset # Load a specific batch config ds = load_dataset("avewright/tabula-pretraining-corpus", name="datagen_001") # Load all configs from… See the full description on the dataset page: https://huggingface.co/datasets/avewright/tabula-pretraining-corpus.tabulartabular-classification1M<n<10M0 likes184 downloads6mo agoHugging Face19aaronfeller /peptideclm-2-pretraining-data PeptideMTR Training Data This repository contains the dataset for the PeptideMTR paper. It is designed for SMILES encoder models trained by masked-language modeling (MLM) and/or multi-target regression (MTR) tasks, focusing on mapping peptide sequences to biochemical properties. Link to the manuscript will be added here when available. Dataset Summary The dataset includes peptide sequences paired with 99 RDKit-derived descriptors representing various physicochemical… See the full description on the dataset page: https://huggingface.co/datasets/aaronfeller/peptideclm-2-pretraining-data.tabular100M<n<1B0 likes173 downloads9mo agoHugging Face20lapa-llm /pretraining-high-quality Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness of the text lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.tabulartext-generation10M<n<100M0 likes139 downloads10mo agoHugging Face21superAVTR /wikipedia_en_512_for_pretraining Cleaned Wikipedia 512 Pretraining Dataset Dataset Description This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining. The original dataset contains English Wikipedia text prepared for language-model pretraining. This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text. Source Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.tabular1M<n<10M0 likes136 downloads5d agoHugging Face22simple-pretraining /bookcorpusopen_with_ids_chunked Dataset Card for "bookcorpusopen_with_ids_chunked" More Information needed tabular10M<n<100M0 likes104 downloads3y agoHugging Face23lapa-llm /pretraining-lower-quality Dataset Card for Lapa Pretraining Lower Quality Dataset Dataset Description Dataset Summary This dataset is a high quality (but lower quality than https://huggingface.co/datasets/lapa-llm/pretraining-high-qualit) subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality.tabulartext-generation10M<n<100M0 likes103 downloads10mo agoHugging Face24ZhuofengLi /pretraining-pretokenized-smollm3 SmolLM3 Pretokenized Pretraining Sources Datatrove/Nanotron tokenized-byte versions of three public pretraining sources: fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT finemath-4plus: HuggingFaceTB/finemath, finemath-4plus stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.tabularn<1K0 likes98 downloads2mo agoHugging Face25gashingriver5963 /DCLM-pretraining-datasettabular1K<n<10K0 likes72 downloads2mo agoHugging Face26odoma /gotriple-pretraining-dataset GoTriple Pretraining dataset Summary The GoTriple Pre-training Dataset is a multilingual corpus built from open-access research artefacts harvested via the GoTriple platform. It focuses on Social Sciences and Humanities (SSH) content, addressing their limited presence in standard LLM pre-training corpora. Current release includes History, Sociology, Environmental Sciences, Psychology and Geography texts (~23.14B tokens). Intended Use Continuous… See the full description on the dataset page: https://huggingface.co/datasets/odoma/gotriple-pretraining-dataset.tabular100K<n<1M3 likes70 downloads8mo agoHugging Face27Eugleo /pretraining-priors-political-partisan pretraining-priors political partisan answers For every one of the 6,820 IdeoINST questions (Chen et al., EMNLP 2024, arXiv:2402.11725; the prompt set Eugleo/pretraining-priors-political-eval), one claude-sonnet-5 call wrote a very strongly and clearly left-leaning answer and a very strongly and clearly right-leaning answer (about three sentences each, first person, no partisan self-label), and tagged the question with one 0/1 flag per domain. A flag cat_<domain> is 1 when a… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-political-partisan.tabular1K<n<10K0 likes60 downloads12d agoHugging Face28haydn-jones /ORD-Pretraining-v5tabular1M<n<10M0 likes56 downloads7mo agoHugging Face29ChengsenWang /GenoJEPA-Pretraining GenoJEPA-Pretraining This dataset provides the pre-training resources used for GenoJEPA, a genomic representation learning framework based on joint-embedding predictive architecture. GenoJEPA learns semantic representations of DNA sequences by shifting the optimization target from nucleotide-level reconstruction to latent-space semantic alignment. The pre-training data is used to construct global and local sequence views for self-supervised genomic representation learning.… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/GenoJEPA-Pretraining.tabularn<1K0 likes47 downloads24d agoHugging Face30kojikojiprg /ai-theories-corpus-en-pretraining ai-theories en コーパス ai-theories(https://github.com/kojikojiprg/ai-theories)プロジェクトの成果物。 006(小型 GPT の事前学習)・008 の英語条件用のコーパス。 ライセンスについての注記 このデータセット自体のライセンスは クリエイティブ・コモンズ 表示-継承 4.0 国際 (CC BY-SA 4.0) である。ai-theoriesの他の成果物(トークナイザなど)は cc-by-nc-4.0を採用しているが、本データセットの内容はフリー百科事典 『ウィキペディア(Wikipedia)』の本文そのもの(wikitext を平文に変換したのみで、 内容は改変していない)であり、Wikipedia 本文自体のライセンス(CC BY-SA 4.0、 表示・継承の条件)を継承する必要があるため、別のライセンスとしている。 由来… See the full description on the dataset page: https://huggingface.co/datasets/kojikojiprg/ai-theories-corpus-en-pretraining.tabularn<1K0 likes45 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.