datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openhermes-2.5-mig-50k
Openhermes-2.5-MIG-50K
Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Openhermes2.5.
Citation
@article{chen2025mig,
title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space},
author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/openhermes-2.5-mig-50k.tulu-3-deita-50k
Tulu-3-DEITA-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using DEITA.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-deita-50k.tulu-3-zip-50k
Tulu-3-ZIP-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K diverse SFT data sampled from Tulu3 using ZIP.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-zip-50k.tulu-3-mig-50k
Tulu-3-MIG-50K
Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-mig-50k.tulu-3-random-50k
Tulu-3-Random-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K SFT data randomly sampled from Tulu3.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-random-50k.tulu-3-qdit-50k
Tulu-3-QDIT-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using QDIT.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-qdit-50k.tulu-3-car-50k
Tulu-3-CaR-50K
Project | Github | Paper | HuggingFace's collection
This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3 using CaR.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-car-50k.deita-sota-mig-6k
Deita-Sota-MIG-6K
Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Deita-Sota-Pool.
Citation
@article{chen2025mig,
title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space},
author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/deita-sota-mig-6k.tulu-3-ifd-50k
Tulu-3-IFD-50K
Project | Github | Paper | HuggingFace's collection
This dataset includes 50K high-quality SFT data sampled from Tulu3 using IFD.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19
6.71
6.64… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-ifd-50k.tulu-3-instag-50k
Tulu-3-InsTag-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using InsTag.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-instag-50k.
