CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aslawliet /math-pretraining-corpustext10M<n<100M4 likes2k downloads2y agoHugging Face02fineinstructions-pretraining /nemotron_qa_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes715 downloads8mo agoHugging Face03fineinstructions-pretraining /nemotron_actual_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes488 downloads8mo agoHugging Face04fineinstructions-pretraining /nemotron_fineinstructions_1T_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B3 likes478 downloads8mo agoHugging Face05fineinstructions-pretraining /nemotron_synthetic_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes348 downloads8mo agoHugging Face06Linly-AI /Chinese-pretraining-datasetData source: https://github.com/CVI-SZU/Linly/wiki/Linly-OpenLLaMA text10M<n<100M44 likes300 downloads3y agoHugging Face07Johnblick187 /tweaktron-pretraining-data-2text1M<n<10M0 likes241 downloads2mo agoHugging Face08visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes209 downloads9mo agoHugging Face09fineinstructions-pretraining /nemotron_wrap_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes184 downloads8mo agoHugging Face10FreedomIntelligence /HuatuoGPT2-Pretraining-Instruction HuatuoGPT2-Pretraining-Instruction-5200K Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT. This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible. Data Volume The following table details the volume and distribution of pre-training data for HuatuoGPT2: Data Source Data Volume Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.textquestion-answering1M<n<10M13 likes164 downloads2y agoHugging Face11superAVTR /wikipedia_en_512_for_pretraining Cleaned Wikipedia 512 Pretraining Dataset Dataset Description This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining. The original dataset contains English Wikipedia text prepared for language-model pretraining. This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text. Source Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.tabular1M<n<10M0 likes143 downloads7d agoHugging Face12windprak /steuerllm_pretraining_dataset SteuerLLM Pretraining Dataset Project page | Paper | GitHub Pretraining Dataset for German Tax Law filtered from FineWeb. This dataset was used for the continual pretraining stage of SteuerLLM, a specialized large language model for German tax law analysis. Dataset Description The SteuerLLM pretraining dataset is a domain-specific subset filtered from large-scale web corpora. It focuses on identifying and extracting tax-related content from German web data to adapt… See the full description on the dataset page: https://huggingface.co/datasets/windprak/steuerllm_pretraining_dataset.texttext-generation1M<n<10M1 likes118 downloads7mo agoHugging Face13fineinstructions-pretraining /ipt_actual_all_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M0 likes111 downloads8mo agoHugging Face14fineinstructions-pretraining /ipt_fineinstructions_all_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes99 downloads8mo agoHugging Face15fineinstructions-pretraining /ipt_synthetic_all_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes91 downloads8mo agoHugging Face16fineinstructions-pretraining /ipt_fineinstructions_all_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes76 downloads8mo agoHugging Face17gashingriver5963 /DCLM-pretraining-datasettabular1K<n<10K0 likes72 downloads2mo agoHugging Face18continual-pretraining /japanese-deduptext100M<n<1B2 likes49 downloads2y agoHugging Face19huu-ontocord /test-codeact-pretrainingtext1K<n<10K0 likes48 downloads5mo agoHugging Face20kojikojiprg /ai-theories-corpus-en-pretraining ai-theories en コーパス ai-theories(https://github.com/kojikojiprg/ai-theories)プロジェクトの成果物。 006(小型 GPT の事前学習)・008 の英語条件用のコーパス。 ライセンスについての注記 このデータセット自体のライセンスは クリエイティブ・コモンズ 表示-継承 4.0 国際 (CC BY-SA 4.0) である。ai-theoriesの他の成果物(トークナイザなど)は cc-by-nc-4.0を採用しているが、本データセットの内容はフリー百科事典 『ウィキペディア(Wikipedia)』の本文そのもの(wikitext を平文に変換したのみで、 内容は改変していない)であり、Wikipedia 本文自体のライセンス(CC BY-SA 4.0、 表示・継承の条件)を継承する必要があるため、別のライセンスとしている。 由来… See the full description on the dataset page: https://huggingface.co/datasets/kojikojiprg/ai-theories-corpus-en-pretraining.tabularn<1K0 likes45 downloads11d agoHugging Face21open-llm-leaderboard /FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face22open-llm-leaderboard /FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face23fineinstructions-pretraining /ipt_fineinstructions_all_judged_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes39 downloads8mo agoHugging Face24shuyuej /English-Pretraining-Datasettext10K<n<100K3 likes33 downloads2y agoHugging Face25Ashmal /Arabic_Pretraining_10Ktext10K<n<100K0 likes30 downloads2y agoHugging Face26FishCaduceus /FishCaduceus-Pretraining-1024 FishCaduceus Pretraining Dataset 1024 Dataset description This dataset contains fixed-length 1,024-bp genomic sequence windows prepared for pretraining FishCaduceus DNA language models. It uses single-nucleotide tokenization and is intended for masked language modeling. Dataset summary The dataset contains 2,709,306 JSON Lines records. Each record contains genomic coordinates and a seq field. The six files were audited by streaming line-by-line… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Pretraining-1024.tabular1M<n<10M0 likes30 downloads3mo agoHugging Face27Papajams /repro-gram-modular-pretraining-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes30 downloads1mo agoHugging Face28living-models /Botanic1-pretraininggated Botanic1 pretraining data This dataset accompanies the Botanic1 technical report. It preserves the genomic sequence shards used to pretrain the Botanic1 models, including the original sequence case, augmentation margins, train/test separation, source manifests, and assembly metadata. Release contents The initial release contains the main 8 kbp corpus used by Botanic1-S, M, L, and XL. The context extension corpora at 16, 32, 64, and 128 kbp are planned for this… See the full description on the dataset page: https://huggingface.co/datasets/living-models/Botanic1-pretraining.textfill-mask1M<n<10M0 likes30 downloads14d agoHugging Face29michaelchenkj /150M-0.025x-DCLM-pretraining-datasettabular1K<n<10K0 likes27 downloads9mo agoHugging Face30zhqwqwq /NCPL-Pretraining-LogsPretraining logs collected from: Marin Project: https://github.com/marin-community/marin Step Law Project: https://github.com/step-law/steplaw Each example corresponds to one training run, including the training configuration and performance metrics (C4-en evaluation loss for Marin, and smoothed pretraining loss for StepLaw). Load the dataset from datasets import load_dataset marin = load_dataset("zhqwqwq/NCPL-Pretraining-Logs", "marin", split="train") steplaw =… See the full description on the dataset page: https://huggingface.co/datasets/zhqwqwq/NCPL-Pretraining-Logs.tabular1K<n<10K0 likes26 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.