datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flanThis is a repreprocessed version of the FLAN dataset with any updates that have been made to the FLAN datasets since the release of the original FLAN. The script is available here.
Tasks:
{'aeslc_10templates',
'ag_news_subset_10templates',
'anli_r1_10templates',
'anli_r2_10templates',
'anli_r3_10templates',
'arc_challenge_10templates',
'arc_easy_10templates',
'bool_q_10templates',
'cb_10templates',
'cnn_dailymail_10templates',
'cola_10templates',
'common_gen_10templates'… See the full description on the dataset page: https://huggingface.co/datasets/Muennighoff/flan.flan_v2
Dataset Card for Flan V2
Dataset Summary
This is a processed version of the Flan V2 dataset.
I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing.
The authors of the Flan Collection recommend experimenting with different mixing ratio's of tasks to get optimal results downstream.
Setup Instructions
Here are the steps I followed to get everything working:
Build AESLC and WinoGrande datasets… See the full description on the dataset page: https://huggingface.co/datasets/SirNeural/flan_v2.flanv2flanv2
Fork of SirNeural/flan_v2
just in case it gets deleted.
Dataset Card for Flan V2
Dataset Summary
This is a processed version of the Flan V2 dataset.
I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing.
The authors of the Flan Collection recommend experimenting with different mixing ratio's of tasks to get optimal results downstream.
This current version I've processed is missing a few… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/flanv2.flan2021-full
Task Name
FLAN-2021 -> 70
{
"ag_news_subset": 108497,
"ai2_arc/ARC-Challenge": 829,
"ai2_arc/ARC-Easy": 1927,
"aeslc": 13187,
"anli/r1": 15361,
"anli/r2": 41133,
"anli/r3": 91048,
"bool_q": 8343,
"cnn_dailymail": 259607,
"coqa": 6456,
"cosmos_qa": 22996,
"definite_pronoun_resolution": 1079,
"drop": 70045,
"fix_punct": 25690,
"gem/common_gen": 60936,
"gem/dart": 56724,
"gem/e2e_nlg": 30337,
"gem/web_nlg_en": 31899… See the full description on the dataset page: https://huggingface.co/datasets/aslawliet/flan2021-full.flan-embed-test2Commercial-Flan-Collection-Chain-Of-Thoughtk3-sft-cc0-flan
Dataset Card for K3 SFT CC0 FLAN
844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain
FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort
views; adaptive is the recommended default for quality-conscious SFT mixing.
Dataset Details
Curated by: Training Datasmith
Teacher: kimi-k3 via deltafin (local inference)
Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.FLAN-Small
FLAN-Small
This repository is a reduced version of the data provided by the hardwork of: https://huggingface.co/datasets/imone/OpenOrca_FLAN.
FLAN-Small amounts to ~10m examples sampled to approximately hold to the FLAN's final "submix" of:
{
'flan': 0.4,
't0': 0.32,
'niv2': 0.20,
'cot': 0.05,
'dialog': 0.03
}
Since the cot data is rather small -- this was sampled with replacement; consequently there are some duplicates.
Some token length… See the full description on the dataset page: https://huggingface.co/datasets/BadDepartment/FLAN-Small.flan_labeledCommercial-Flan-Collection-Dialogflan1m-alpaca-uncensoredaxolotl was giving me issues with dolphin. please give all credit and support to https://huggingface.co/ehartford!
sapient-synth-flan-niv2-fsopt-data-task902-deceptive-opinion-spam-classification
sapient-synth-flan-niv2-fsopt-data-task902-deceptive-opinion-spam-classification
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 2935
Task: synthetic anonymous instruction replacement
Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task902-deceptive-opinion-spam-classification.sapient-synth-flan-niv2-fsopt-data-task635-allegro-reviews-answer-generation
sapient-synth-flan-niv2-fsopt-data-task635-allegro-reviews-answer-generation
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 3186
Task: synthetic anonymous instruction replacement
Generation
Rows… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task635-allegro-reviews-answer-generation.flangoogle__flan-t5-large-details
Dataset Card for Evaluation run of google/flan-t5-large
Dataset automatically created during the evaluation run of model google/flan-t5-large
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-large-details.subsampled_flan_v2FlanCoT_100k_seq
Dataset Card for FlanCoT-100k-Seq
It is a converted 100k instances in CoT subset of FLAN (Apache 2.0): using Seq-Instruct method from the paper SIT: Fine-tuning Large Language Models with Sequential Instructions
It includes massive sequential instructions which contains subtask more than one in each instructions.
sapient-synth-flan-niv2-fsopt-data-task1377-newscomm-translation
sapient-synth-flan-niv2-fsopt-data-task1377-newscomm-translation
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 3491
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task1377-newscomm-translation.google__flan-t5-small-details
Dataset Card for Evaluation run of google/flan-t5-small
Dataset automatically created during the evaluation run of model google/flan-t5-small
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-small-details.google__flan-t5-base-details
Dataset Card for Evaluation run of google/flan-t5-base
Dataset automatically created during the evaluation run of model google/flan-t5-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-base-details.google__flan-ul2-details
Dataset Card for Evaluation run of google/flan-ul2
Dataset automatically created during the evaluation run of model google/flan-ul2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-ul2-details.BlendNet
📚 BlendNet
The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples.
For more details, please visit our GitHub repository or refer to our arXiv paper.
📖 Citation
@misc{du2024blenderllmtraininglargelanguage,
title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement},
author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/FlandreScarlet123/BlendNet.google__flan-t5-xl-details
Dataset Card for Evaluation run of google/flan-t5-xl
Dataset automatically created during the evaluation run of model google/flan-t5-xl
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-xl-details.flan_t5_qnaflan_t5_qna
flan2021_submix_original_subset_v1google__flan-t5-xxl-details
Dataset Card for Evaluation run of google/flan-t5-xxl
Dataset automatically created during the evaluation run of model google/flan-t5-xxl
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-xxl-details.GPT-5.5-Long-Reasoninghi-kn-FLANopen-instruct-flan-synthetic-finetuning
