datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tulu-3-pool-annotated
Tulu-3-Pool-Annotated
Project | Github | Paper | HuggingFace's collection
Annotated tulu-3-sft-mixture. Used as a data pool in MIG. The annotations include #InsTag tags, DEITA scores, and CaR scores.
Dataset Details
Tulu3 Dataset Sources
Repository: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture
Paper [optional]: Tulu 3: Pushing Frontiers in Open Language Model Post-Training
MIG Dataset Sources
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-pool-annotated.openhermes-2.5-mig-50k
Openhermes-2.5-MIG-50K
Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Openhermes2.5.
Citation
@article{chen2025mig,
title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space},
author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/openhermes-2.5-mig-50k.tulu-3-deita-50k
Tulu-3-DEITA-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using DEITA.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-deita-50k.tulu-3-zip-50k
Tulu-3-ZIP-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K diverse SFT data sampled from Tulu3 using ZIP.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-zip-50k.tulu-3-mig-50k
Tulu-3-MIG-50K
Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-mig-50k.tulu-3-random-50k
Tulu-3-Random-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K SFT data randomly sampled from Tulu3.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-random-50k.tulu-3-qdit-50k
Tulu-3-QDIT-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using QDIT.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-qdit-50k.openhermes-2.5-pool-annotated
Openhermes-2.5-Pool-Annotated
Annotated Openhermes-2.5. Used as a data pool in MIG. The annotations include #InsTag tags, DEITA scores, and CaR scores.
Dataset Details
Dataset Sources
Repository: https://huggingface.co/datasets/teknium/OpenHermes-2.5
Citation
BibTeX:
@misc{OpenHermes 2.5,
title = {OpenHermes 2.5: An Open Dataset of Synthetic Data for Generalist LLM Assistants},
author = {Teknium},
year = {2023},
publisher… See the full description on the dataset page: https://huggingface.co/datasets/xsample/openhermes-2.5-pool-annotated.CF-Worldtulu-3-car-50k
Tulu-3-CaR-50K
Project | Github | Paper | HuggingFace's collection
This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3 using CaR.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-car-50k.deita-sota-pool-annotated
Deita-Sota-Pool-Annotated
Annotated Deia-Sota-Pool. Used as a data pool in MIG. The annotations include #InsTag tags, DEITA scores, and CaR scores.
Dataset Details
Dataset Sources
Repository: https://huggingface.co/datasets/AndrewZeng/deita_sota_pool
Citation
BibTeX:
deita-sota-mig-6k
Deita-Sota-MIG-6K
Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Deita-Sota-Pool.
Citation
@article{chen2025mig,
title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space},
author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/deita-sota-mig-6k.tulu-3-ifd-50k
Tulu-3-IFD-50K
Project | Github | Paper | HuggingFace's collection
This dataset includes 50K high-quality SFT data sampled from Tulu3 using IFD.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19
6.71
6.64… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-ifd-50k.tulu-3-instag-50k
Tulu-3-InsTag-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using InsTag.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-instag-50k.
