datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia_en_512_for_pretraining
Cleaned Wikipedia 512 Pretraining Dataset
Dataset Description
This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining.
The original dataset contains English Wikipedia text prepared for language-model pretraining.
This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text.
Source
Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.DCLM-pretraining-datasetai-theories-corpus-en-pretraining
ai-theories en コーパス
ai-theories(https://github.com/kojikojiprg/ai-theories)プロジェクトの成果物。
006(小型 GPT の事前学習)・008 の英語条件用のコーパス。
ライセンスについての注記
このデータセット自体のライセンスは クリエイティブ・コモンズ 表示-継承 4.0 国際
(CC BY-SA 4.0) である。ai-theoriesの他の成果物(トークナイザなど)は
cc-by-nc-4.0を採用しているが、本データセットの内容はフリー百科事典
『ウィキペディア(Wikipedia)』の本文そのもの(wikitext を平文に変換したのみで、
内容は改変していない)であり、Wikipedia 本文自体のライセンス(CC BY-SA 4.0、
表示・継承の条件)を継承する必要があるため、別のライセンスとしている。
由来… See the full description on the dataset page: https://huggingface.co/datasets/kojikojiprg/ai-theories-corpus-en-pretraining.FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.FishCaduceus-Pretraining-1024
FishCaduceus Pretraining Dataset 1024
Dataset description
This dataset contains fixed-length 1,024-bp genomic sequence windows prepared for pretraining FishCaduceus DNA language models. It uses single-nucleotide tokenization and is intended for masked language modeling.
Dataset summary
The dataset contains 2,709,306 JSON Lines records. Each record contains genomic coordinates and a seq field. The six files were audited by streaming line-by-line… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Pretraining-1024.150M-0.025x-DCLM-pretraining-datasetNCPL-Pretraining-LogsPretraining logs collected from:
Marin Project: https://github.com/marin-community/marin
Step Law Project: https://github.com/step-law/steplaw
Each example corresponds to one training run, including the training configuration and performance metrics (C4-en evaluation loss for Marin, and smoothed pretraining loss for StepLaw).
Load the dataset
from datasets import load_dataset
marin = load_dataset("zhqwqwq/NCPL-Pretraining-Logs", "marin", split="train")
steplaw =… See the full description on the dataset page: https://huggingface.co/datasets/zhqwqwq/NCPL-Pretraining-Logs.FishCaduceus-Pretraining-512
FishCaduceus Pretraining Dataset 512
Dataset description
This dataset contains fixed-length 512-bp genomic sequence windows prepared for pretraining FishCaduceus DNA language models. It uses single-nucleotide tokenization and is intended for masked language modeling.
Dataset summary
The dataset contains 6,087,221 JSON Lines records. Each record contains genomic coordinates and a seq field. The six files were audited by streaming line-by-line… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Pretraining-512.FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.BatatasLM-SP5-Pretraining-Data
BatatasLM-SP5-Pretraining-Data
Dataset Description
A 4096 bp DNA pretraining corpus constructed from five sweetpotato genomes or accessions.
Used By
BatatasLM-SP5-20L_pretrain
BatatasLM-SP5-28L_pretrain
Verified Summary Statistics
Split
Windows
Base pairs
train
340479
1394601984
valid
42077
172347392
test
42524
174178304
The values above are read from the supplied JSONL summary TSV. No unsupported sample counts… See the full description on the dataset page: https://huggingface.co/datasets/BatatasLM/BatatasLM-SP5-Pretraining-Data.add-sub-pre-trainingFlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.DCLM-pretraining-datasetFlofloB__test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.FlofloB__83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.BatatasLM-MS12-Pretraining-Data
BatatasLM-MS12-Pretraining-Data
Dataset Description
A 4096 bp DNA pretraining corpus constructed from a 12-species multi-species plant genome collection.
Used By
BatatasLM-MS12-28L_pretrain
Verified Summary Statistics
Split
Windows
Base pairs
train
385624
1579515904
valid
48395
198225920
test
48103
197029888
The values above are read from the supplied JSONL summary TSV. The supplied genome split summary TSV is… See the full description on the dataset page: https://huggingface.co/datasets/BatatasLM/BatatasLM-MS12-Pretraining-Data.
