CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aslawliet /math-pretraining-corpustext10M<n<100M4 likes2k downloads2y agoHugging Face02fineinstructions-pretraining /nemotron_qa_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes769 downloads8mo agoHugging Face03fineinstructions-pretraining /nemotron_actual_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes532 downloads8mo agoHugging Face04fineinstructions-pretraining /nemotron_fineinstructions_1T_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B3 likes518 downloads8mo agoHugging Face05fineinstructions-pretraining /nemotron_synthetic_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes388 downloads8mo agoHugging Face06Linly-AI /Chinese-pretraining-datasetData source: https://github.com/CVI-SZU/Linly/wiki/Linly-OpenLLaMA text10M<n<100M44 likes275 downloads3y agoHugging Face07Johnblick187 /tweaktron-pretraining-data-2text1M<n<10M0 likes241 downloads2mo agoHugging Face08visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes205 downloads9mo agoHugging Face09anthonyyazdaniml /gliner-biomed-pre-training GLiNER-BioMed pre-training dataset This dataset, used for the pre-training stage of the GLiNER-BioMed models, was introduced in the paper GLiNER-BioMed: a suite of efficient models for open biomedical named entity recognition. Citation If you use the GLiNER-BioMed models or datasets in your work, please cite: @article{yazdani2026gliner, author = {Yazdani, Anthony and Stepanov, Ihor and Teodoro, Douglas}, title = {{GLiNER-BioMed}: a suite of efficient… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-pre-training.text10K<n<100K0 likes201 downloads2mo agoHugging Face10fineinstructions-pretraining /nemotron_wrap_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes196 downloads8mo agoHugging Face11FreedomIntelligence /HuatuoGPT2-Pretraining-Instruction HuatuoGPT2-Pretraining-Instruction-5200K Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT. This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible. Data Volume The following table details the volume and distribution of pre-training data for HuatuoGPT2: Data Source Data Volume Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.textquestion-answering1M<n<10M13 likes160 downloads2y agoHugging Face12superAVTR /wikipedia_en_512_for_pretraining Cleaned Wikipedia 512 Pretraining Dataset Dataset Description This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining. The original dataset contains English Wikipedia text prepared for language-model pretraining. This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text. Source Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.tabular1M<n<10M0 likes136 downloads6d agoHugging Face13windprak /steuerllm_pretraining_dataset SteuerLLM Pretraining Dataset Project page | Paper | GitHub Pretraining Dataset for German Tax Law filtered from FineWeb. This dataset was used for the continual pretraining stage of SteuerLLM, a specialized large language model for German tax law analysis. Dataset Description The SteuerLLM pretraining dataset is a domain-specific subset filtered from large-scale web corpora. It focuses on identifying and extracting tax-related content from German web data to adapt… See the full description on the dataset page: https://huggingface.co/datasets/windprak/steuerllm_pretraining_dataset.texttext-generation1M<n<10M1 likes121 downloads7mo agoHugging Face14fineinstructions-pretraining /ipt_actual_all_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M0 likes117 downloads8mo agoHugging Face15fineinstructions-pretraining /ipt_fineinstructions_all_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes98 downloads8mo agoHugging Face16fineinstructions-pretraining /ipt_synthetic_all_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes98 downloads8mo agoHugging Face17fineinstructions-pretraining /ipt_fineinstructions_all_judged_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes95 downloads8mo agoHugging Face18fineinstructions-pretraining /ipt_fineinstructions_all_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes76 downloads8mo agoHugging Face19gashingriver5963 /DCLM-pretraining-datasettabular1K<n<10K0 likes72 downloads2mo agoHugging Face20continual-pretraining /japanese-deduptext100M<n<1B2 likes49 downloads2y agoHugging Face21huu-ontocord /test-codeact-pretrainingtext1K<n<10K0 likes45 downloads5mo agoHugging Face22kojikojiprg /ai-theories-corpus-en-pretraining ai-theories en コーパス ai-theories(https://github.com/kojikojiprg/ai-theories)プロジェクトの成果物。 006(小型 GPT の事前学習)・008 の英語条件用のコーパス。 ライセンスについての注記 このデータセット自体のライセンスは クリエイティブ・コモンズ 表示-継承 4.0 国際 (CC BY-SA 4.0) である。ai-theoriesの他の成果物(トークナイザなど)は cc-by-nc-4.0を採用しているが、本データセットの内容はフリー百科事典 『ウィキペディア(Wikipedia)』の本文そのもの(wikitext を平文に変換したのみで、 内容は改変していない)であり、Wikipedia 本文自体のライセンス(CC BY-SA 4.0、 表示・継承の条件)を継承する必要があるため、別のライセンスとしている。 由来… See the full description on the dataset page: https://huggingface.co/datasets/kojikojiprg/ai-theories-corpus-en-pretraining.tabularn<1K0 likes45 downloads10d agoHugging Face23open-llm-leaderboard /FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face24open-llm-leaderboard /FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face25TharunSivamani /pretraining_subsets_corpustext1M<n<10M0 likes39 downloads9mo agoHugging Face26shuyuej /English-Pretraining-Datasettext10K<n<100K3 likes33 downloads2y agoHugging Face27Papajams /repro-gram-modular-pretraining-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes33 downloads1mo agoHugging Face28Ashmal /Arabic_Pretraining_10Ktext10K<n<100K0 likes30 downloads2y agoHugging Face29living-models /Botanic1-pretraininggated Botanic1 pretraining data This dataset accompanies the Botanic1 technical report. It preserves the genomic sequence shards used to pretrain the Botanic1 models, including the original sequence case, augmentation margins, train/test separation, source manifests, and assembly metadata. Release contents The initial release contains the main 8 kbp corpus used by Botanic1-S, M, L, and XL. The context extension corpora at 16, 32, 64, and 128 kbp are planned for this… See the full description on the dataset page: https://huggingface.co/datasets/living-models/Botanic1-pretraining.textfill-mask1M<n<10M0 likes30 downloads13d agoHugging Face30michaelchenkj /150M-0.025x-DCLM-pretraining-datasettabular1K<n<10K0 likes27 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.