CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aslawliet /math-pretraining-corpustext10M<n<100M4 likes1.9k downloads2y agoHugging Face02fineinstructions-pretraining /nemotron_qa_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes648 downloads8mo agoHugging Face03fineinstructions-pretraining /nemotron_actual_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes474 downloads8mo agoHugging Face04fineinstructions-pretraining /nemotron_fineinstructions_1T_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B3 likes431 downloads8mo agoHugging Face05fineinstructions-pretraining /nemotron_synthetic_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes324 downloads8mo agoHugging Face06Linly-AI /Chinese-pretraining-datasetData source: https://github.com/CVI-SZU/Linly/wiki/Linly-OpenLLaMA text10M<n<100M44 likes317 downloads3y agoHugging Face07visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes232 downloads9mo agoHugging Face08FreedomIntelligence /HuatuoGPT2-Pretraining-Instruction HuatuoGPT2-Pretraining-Instruction-5200K Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT. This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible. Data Volume The following table details the volume and distribution of pre-training data for HuatuoGPT2: Data Source Data Volume Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.textquestion-answering1M<n<10M13 likes178 downloads2y agoHugging Face09fineinstructions-pretraining /nemotron_wrap_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes177 downloads8mo agoHugging Face10superAVTR /wikipedia_en_512_for_pretraining Cleaned Wikipedia 512 Pretraining Dataset Dataset Description This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining. The original dataset contains English Wikipedia text prepared for language-model pretraining. This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text. Source Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.tabular1M<n<10M0 likes157 downloads9d agoHugging Face11windprak /steuerllm_pretraining_dataset SteuerLLM Pretraining Dataset Project page | Paper | GitHub Pretraining Dataset for German Tax Law filtered from FineWeb. This dataset was used for the continual pretraining stage of SteuerLLM, a specialized large language model for German tax law analysis. Dataset Description The SteuerLLM pretraining dataset is a domain-specific subset filtered from large-scale web corpora. It focuses on identifying and extracting tax-related content from German web data to adapt… See the full description on the dataset page: https://huggingface.co/datasets/windprak/steuerllm_pretraining_dataset.texttext-generation1M<n<10M1 likes113 downloads8mo agoHugging Face12fineinstructions-pretraining /ipt_actual_all_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M0 likes103 downloads8mo agoHugging Face13fineinstructions-pretraining /ipt_synthetic_all_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes91 downloads8mo agoHugging Face14fineinstructions-pretraining /ipt_fineinstructions_all_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes85 downloads8mo agoHugging Face15gashingriver5963 /DCLM-pretraining-datasettabular1K<n<10K0 likes70 downloads2mo agoHugging Face16fineinstructions-pretraining /ipt_fineinstructions_all_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes67 downloads8mo agoHugging Face17kojikojiprg /ai-theories-corpus-en-pretraining ai-theories en コーパス ai-theories(https://github.com/kojikojiprg/ai-theories)プロジェクトの成果物。 006(小型 GPT の事前学習)・008 の英語条件用のコーパス。 ライセンスについての注記 このデータセット自体のライセンスは クリエイティブ・コモンズ 表示-継承 4.0 国際 (CC BY-SA 4.0) である。ai-theoriesの他の成果物(トークナイザなど)は cc-by-nc-4.0を採用しているが、本データセットの内容はフリー百科事典 『ウィキペディア(Wikipedia)』の本文そのもの(wikitext を平文に変換したのみで、 内容は改変していない)であり、Wikipedia 本文自体のライセンス(CC BY-SA 4.0、 表示・継承の条件)を継承する必要があるため、別のライセンスとしている。 由来… See the full description on the dataset page: https://huggingface.co/datasets/kojikojiprg/ai-theories-corpus-en-pretraining.tabularn<1K0 likes55 downloads1d agoHugging Face18continual-pretraining /japanese-deduptext100M<n<1B2 likes49 downloads2y agoHugging Face19open-llm-leaderboard /FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face20fineinstructions-pretraining /ipt_fineinstructions_all_judged_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text10M<n<100M1 likes41 downloads8mo agoHugging Face21open-llm-leaderboard /FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face22Papajams /repro-gram-modular-pretraining-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes34 downloads1mo agoHugging Face23shuyuej /English-Pretraining-Datasettext10K<n<100K3 likes33 downloads2y agoHugging Face24huu-ontocord /test-codeact-pretrainingtext1K<n<10K0 likes32 downloads5mo agoHugging Face25FishCaduceus /FishCaduceus-Pretraining-1024 FishCaduceus Pretraining Dataset 1024 Dataset description This dataset contains fixed-length 1,024-bp genomic sequence windows prepared for pretraining FishCaduceus DNA language models. It uses single-nucleotide tokenization and is intended for masked language modeling. Dataset summary The dataset contains 2,709,306 JSON Lines records. Each record contains genomic coordinates and a seq field. The six files were audited by streaming line-by-line… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Pretraining-1024.tabular1M<n<10M0 likes31 downloads3mo agoHugging Face26Ashmal /Arabic_Pretraining_10Ktext10K<n<100K0 likes30 downloads2y agoHugging Face27living-models /Botanic1-pretraininggated Botanic1 pretraining data This dataset accompanies the Botanic1 technical report. It preserves the genomic sequence shards used to pretrain the Botanic1 models, including the original sequence case, augmentation margins, train/test separation, source manifests, and assembly metadata. Release contents The initial release contains the main 8 kbp corpus used by Botanic1-S, M, L, and XL. The context extension corpora at 16, 32, 64, and 128 kbp are planned for this… See the full description on the dataset page: https://huggingface.co/datasets/living-models/Botanic1-pretraining.textfill-mask1M<n<10M0 likes30 downloads16d agoHugging Face28michaelchenkj /150M-0.025x-DCLM-pretraining-datasettabular1K<n<10K0 likes27 downloads9mo agoHugging Face29zhqwqwq /NCPL-Pretraining-LogsPretraining logs collected from: Marin Project: https://github.com/marin-community/marin Step Law Project: https://github.com/step-law/steplaw Each example corresponds to one training run, including the training configuration and performance metrics (C4-en evaluation loss for Marin, and smoothed pretraining loss for StepLaw). Load the dataset from datasets import load_dataset marin = load_dataset("zhqwqwq/NCPL-Pretraining-Logs", "marin", split="train") steplaw =… See the full description on the dataset page: https://huggingface.co/datasets/zhqwqwq/NCPL-Pretraining-Logs.tabular1K<n<10K0 likes24 downloads7mo agoHugging Face30shuyuej /French-Pretraining-Datasettext1K<n<10K1 likes19 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.