CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01applied-ai-018 /pretraining_v1-omega_bookstabular100M<n<1B25 likes377k downloads2y agoHugging Face02projectkaira /Pretraining-V1 Indic TTS Unified v1 A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio. All audio is resampled to 24 kHz mono. Every row follows an identical schema regardless of source… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/Pretraining-V1.audiotext-to-speech10M<n<100M0 likes13k downloads2mo agoHugging Face03allenai /MolmoAct-Pretraining-Mixture MolmoAct - Pretraining Mixture Data Mixture used for MolmoAct Pretraining. Contains a subset of OXE formulated as Action Reasoning Data along with auxiliary robot data and link to Multimodal Web data. MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Pretraining-Mixture.imagerobotics10M<n<100M14 likes12k downloads1y agoHugging Face04lightonai /multilingual-embeddings-pre-training-curated 📚 Collection | 📝 Multilingual Blog | 📝 English Blog Contrastive Multilingual Pre-Training 2.16B query–document pairs across eight languages, plus cross-lingual pairs mDenseOn | mLateOn | DenseOn | LateOn | PyLate | FastPlaid 🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.texttext-retrieval1B<n<10B5 likes9.4k downloads2mo agoHugging Face05nvidia /Nemotron-Pretraining-Specialized-v1 Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.texttext-generation10M<n<100M87 likes5.7k downloads9mo agoHugging Face06geodesic-research /control-pretraining-datasets-smoke geodesic-research/control-pretraining-datasets-smoke Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/control-pretraining-datasets-smoke", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/control-pretraining-datasets-smoke.text10K<n<100K0 likes4k downloads18d agoHugging Face07nvidia /Nemotron-Pretraining-Code-v2gated Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2.texttext-generation100M<n<1B135 likes3.8k downloads9mo agoHugging Face08nvidia /Nemotron-Pretraining-Code-v1gated Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1.texttext-generation100M<n<1B78 likes3.8k downloads9mo agoHugging Face09HuggingFaceBio /carbon-pretraining-corpus 🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.tabulartext-generation100M<n<1B30 likes3.4k downloads3mo agoHugging Face10nvidia /Nemotron-Pretraining-Specialized-v1.2 Nemotron-Pretraining-Specialized-v1.2 Dataset Description: The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions. Note: These are new datasets, not replacements.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.2.texttext-generation100M<n<1B16 likes2.5k downloads4mo agoHugging Face11nvidia /Nemotron-Pretraining-Specialized-v1.1 Nemotron-Pretraining-Specialized-v1.1 Dataset Description: The Nemotron-Pretraining-Specialized-v1.1 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities in code concepts and algorithms, formal logic, economics, and multiple choice questions. The code concepts dataset is an instance of a general… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.1.texttext-generation10M<n<100M46 likes2.4k downloads7mo agoHugging Face12tingtang2 /the_stack_v2_python_repos_pretraining_dataset_imported_context-datasettext1M<n<10M0 likes2.2k downloads1y agoHugging Face13nvidia /Nemotron-Pretraining-SFT-v1gated Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1.texttext-generation100M<n<1B73 likes1.8k downloads9mo agoHugging Face14JuaAI /ts-icl-pretraining-corpus TS-ICL Pretraining Corpus (community reconstruction) A unified, cleaned reconstruction of the univariate pretraining corpus described in Table 5 of TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning (Le Naour, Nabil & Petralia, EDF R&D; arXiv:2606.05878). The TS-ICL authors did not release their pretraining data pipeline, so this corpus is rebuilt from the named upstream sources (LOTSA, Chronos, and the TempoPFN synthetic generators) and… See the full description on the dataset page: https://huggingface.co/datasets/JuaAI/ts-icl-pretraining-corpus.tabulartime-series-forecasting1M<n<10M0 likes1.7k downloads3mo agoHugging Face15avewright /tabula-pretraining-corpus-v2 Tabula Pretraining Corpus v2 A large-scale synthetic tabular dataset for pretraining transformer-based in-context learning models for tabular data (similar to TabPFN). Overview Metric Value Total rows 272,271,776 Total datasets 10,867 Shards 135 Mean utility AUC 0.851 Format Parquet (float32) Schema Each shard is a Parquet file with a fixed-width schema: feat_0 through feat_63: Float32 feature columns. Unused slots are NaN. target:… See the full description on the dataset page: https://huggingface.co/datasets/avewright/tabula-pretraining-corpus-v2.tabulartabular-classification1B<n<10B0 likes1.7k downloads6mo agoHugging Face16aklein4 /seq2seq-mixed-pretraining-SmolLM2tabular100M<n<1B1 likes1.6k downloads8mo agoHugging Face17vaishali /multitabqa_pretraining Usage import pandas as pd from datasets import load_dataset multitableQA_pretraining = load_dataset("vaishali/multitabqa_pretraining") for sample in multitableQA_pretraining['train']: sql_query = sample['query'] input_table_names = sample["table_names"] input_tables = [pd.read_json(table, orient='split') for table in sample['tables']] answer = pd.read_json(sample['answer'], orient='split') # flattened input/output input_to_model = sample["source"] target =… See the full description on the dataset page: https://huggingface.co/datasets/vaishali/multitabqa_pretraining.texttable-question-answering100K<n<1M1 likes1.3k downloads3y agoHugging Face18simple-pretraining /wikipedia_chunked Dataset Card for "wikipedia_chunked" More Information needed text10M<n<100M2 likes1.3k downloads3y agoHugging Face19nvidia /esm2_uniref_pretraining_data ESM-2 Uniref Pretraining Data Dataset Description: UniRef, or UniProt Reference Clusters, are databases of clustered protein sequences from the UniProt Knowledgebase (UniProtKB) that group similar sequences to reduce redundancy and make data easier to work with for biological research. It offers different levels of clustering (UniRef100, UniRef90, and UniRef50) based on sequence identity, with each cluster containing a representative sequence, a count of member proteins… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/esm2_uniref_pretraining_data.textfill-mask100M<n<1B9 likes1.2k downloads1y agoHugging Face20nvidia /Nemotron-Pretraining-Code-v3 Nemotron-Pretraining-Code-v3 Dataset Description: The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs. The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v3.texttext-generation100M<n<1B73 likes1.2k downloads4mo agoHugging Face21nvidia /Nemotron-Pretraining-Legal-v1 Nemotron-Pretraining-Legal-v1 Dataset Description: The Nemotron-Pretraining-Legal-v1 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets intended to improve the legal capabilities of LLMs. In one ablation, adding these datasets to Nemotron 3 Nano pretraining boosted a proxy LegalBench average accuracy from 64.6 to 74.7. This… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Legal-v1.texttext-generation1M<n<10M25 likes1k downloads4mo agoHugging Face22nomic-ai /nomic-bert-2048-pretraining-data Dataset Card for "bert-pretokenized-2048-wiki-2023" More Information needed 1M<n<10M1 likes931 downloads3y agoHugging Face23nvidia /Nemotron-Pretraining-Dataset-sample Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Dataset-sample.text10K<n<100K69 likes910 downloads9mo agoHugging Face24EleutherAI /deep-ignorance-pretraining-mix Deep Ignorance Model Suite We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-pretraining-mix.tabular100M<n<1B4 likes867 downloads1y agoHugging Face25erenyeager-1 /Nemotron-Pretraining-Specialized-v1 Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Nemotron-Pretraining-Specialized-v1.texttext-generation10M<n<100M0 likes667 downloads1mo agoHugging Face26Bingsu /KcBERT_Pre-Training_Corpus KcBERT Pre-Training Corpus (Korean News Comments) KcBERT beomi/kcbert-base Github KcBERT Repo: https://github.com/Beomi/KcBERTKcBERT is Korean Comments BERT pretrained on this Corpus set.(You can use it via Huggingface's Transformers library!) This Kaggle Dataset contains CLEANED dataset preprocessed with the code below. import re import emoji from soynlp.normalizer import repeat_normalize emojis = ''.join(emoji.UNICODE_EMOJI.keys()) pattern = re.compile(f'[^ .… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/KcBERT_Pre-Training_Corpus.textfill-mask10M<n<100M1 likes648 downloads4y agoHugging Face27angie-chen55 /bert_pretraining_data Dataset Card for "bert_pretraining_data" More Information needed 10M<n<100M0 likes636 downloads3y agoHugging Face28tingtang2 /the_stack_v2_2M_repos_pretraining_dataset_imported_context-datasettext100K<n<1M1 likes552 downloads1y agoHugging Face29AmanPriyanshu /stratified-kmeans-diverse-pretraining-100K-1M Stratified K-Means Diverse Pre-Training Dataset (100K-1M) A carefully balanced subset combining FineWeb-Edu and Proof-Pile-2, featuring embedding-based k-means sampling to ensure diverse representation across educational and mathematical/scientific content at multiple scales. 👥 Follow the Authors Aman Priyanshu Supriti Vijay Overview This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales, combining high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-pretraining-100K-1M.texttext-generation1M<n<10M1 likes527 downloads1y agoHugging Face30nasa-ibm-ai4science /Sombench-pretraining-data SomBench Pre-training Corpus: Multimodal Lunar Tiles Dataset Summary This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining. Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.tabularn<1K0 likes519 downloads16d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.