CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01superAVTR /wikipedia_en_512_for_pretraining Cleaned Wikipedia 512 Pretraining Dataset Dataset Description This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining. The original dataset contains English Wikipedia text prepared for language-model pretraining. This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text. Source Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.tabular1M<n<10M0 likes143 downloads6d agoHugging Face02gashingriver5963 /DCLM-pretraining-datasettabular1K<n<10K0 likes72 downloads2mo agoHugging Face03kojikojiprg /ai-theories-corpus-en-pretraining ai-theories en コーパス ai-theories(https://github.com/kojikojiprg/ai-theories)プロジェクトの成果物。 006(小型 GPT の事前学習)・008 の英語条件用のコーパス。 ライセンスについての注記 このデータセット自体のライセンスは クリエイティブ・コモンズ 表示-継承 4.0 国際 (CC BY-SA 4.0) である。ai-theoriesの他の成果物(トークナイザなど)は cc-by-nc-4.0を採用しているが、本データセットの内容はフリー百科事典 『ウィキペディア(Wikipedia)』の本文そのもの(wikitext を平文に変換したのみで、 内容は改変していない)であり、Wikipedia 本文自体のライセンス(CC BY-SA 4.0、 表示・継承の条件)を継承する必要があるため、別のライセンスとしている。 由来… See the full description on the dataset page: https://huggingface.co/datasets/kojikojiprg/ai-theories-corpus-en-pretraining.tabularn<1K0 likes45 downloads11d agoHugging Face04open-llm-leaderboard /FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face05open-llm-leaderboard /FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face06FishCaduceus /FishCaduceus-Pretraining-1024 FishCaduceus Pretraining Dataset 1024 Dataset description This dataset contains fixed-length 1,024-bp genomic sequence windows prepared for pretraining FishCaduceus DNA language models. It uses single-nucleotide tokenization and is intended for masked language modeling. Dataset summary The dataset contains 2,709,306 JSON Lines records. Each record contains genomic coordinates and a seq field. The six files were audited by streaming line-by-line… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Pretraining-1024.tabular1M<n<10M0 likes30 downloads3mo agoHugging Face07michaelchenkj /150M-0.025x-DCLM-pretraining-datasettabular1K<n<10K0 likes27 downloads9mo agoHugging Face08zhqwqwq /NCPL-Pretraining-LogsPretraining logs collected from: Marin Project: https://github.com/marin-community/marin Step Law Project: https://github.com/step-law/steplaw Each example corresponds to one training run, including the training configuration and performance metrics (C4-en evaluation loss for Marin, and smoothed pretraining loss for StepLaw). Load the dataset from datasets import load_dataset marin = load_dataset("zhqwqwq/NCPL-Pretraining-Logs", "marin", split="train") steplaw =… See the full description on the dataset page: https://huggingface.co/datasets/zhqwqwq/NCPL-Pretraining-Logs.tabular1K<n<10K0 likes26 downloads7mo agoHugging Face09FishCaduceus /FishCaduceus-Pretraining-512 FishCaduceus Pretraining Dataset 512 Dataset description This dataset contains fixed-length 512-bp genomic sequence windows prepared for pretraining FishCaduceus DNA language models. It uses single-nucleotide tokenization and is intended for masked language modeling. Dataset summary The dataset contains 6,087,221 JSON Lines records. Each record contains genomic coordinates and a seq field. The six files were audited by streaming line-by-line… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Pretraining-512.tabular1M<n<10M0 likes18 downloads3mo agoHugging Face10open-llm-leaderboard /FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face11BatatasLM /BatatasLM-SP5-Pretraining-Data BatatasLM-SP5-Pretraining-Data Dataset Description A 4096 bp DNA pretraining corpus constructed from five sweetpotato genomes or accessions. Used By BatatasLM-SP5-20L_pretrain BatatasLM-SP5-28L_pretrain Verified Summary Statistics Split Windows Base pairs train 340479 1394601984 valid 42077 172347392 test 42524 174178304 The values above are read from the supplied JSONL summary TSV. No unsupported sample counts… See the full description on the dataset page: https://huggingface.co/datasets/BatatasLM/BatatasLM-SP5-Pretraining-Data.tabular100K<n<1M0 likes10 downloads2mo agoHugging Face12lugman /add-sub-pre-trainingtabular100K<n<1M0 likes9 downloads2mo agoHugging Face13open-llm-leaderboard /FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face14michaelchenkj /DCLM-pretraining-datasettabular1K<n<10K0 likes7 downloads5mo agoHugging Face15open-llm-leaderboard /FlofloB__test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face16open-llm-leaderboard /FlofloB__83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face17BatatasLM /BatatasLM-MS12-Pretraining-Data BatatasLM-MS12-Pretraining-Data Dataset Description A 4096 bp DNA pretraining corpus constructed from a 12-species multi-species plant genome collection. Used By BatatasLM-MS12-28L_pretrain Verified Summary Statistics Split Windows Base pairs train 385624 1579515904 valid 48395 198225920 test 48103 197029888 The values above are read from the supplied JSONL summary TSV. The supplied genome split summary TSV is… See the full description on the dataset page: https://huggingface.co/datasets/BatatasLM/BatatasLM-MS12-Pretraining-Data.tabular100K<n<1M0 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.