datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
XSafetyThe dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license.
tulu-3-pool-annotated
Tulu-3-Pool-Annotated
Project | Github | Paper | HuggingFace's collection
Annotated tulu-3-sft-mixture. Used as a data pool in MIG. The annotations include #InsTag tags, DEITA scores, and CaR scores.
Dataset Details
Tulu3 Dataset Sources
Repository: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture
Paper [optional]: Tulu 3: Pushing Frontiers in Open Language Model Post-Training
MIG Dataset Sources
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-pool-annotated.openhermes-2.5-mig-50k
Openhermes-2.5-MIG-50K
Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Openhermes2.5.
Citation
@article{chen2025mig,
title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space},
author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/openhermes-2.5-mig-50k.tulu-3-deita-50k
Tulu-3-DEITA-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using DEITA.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-deita-50k.tulu-3-zip-50k
Tulu-3-ZIP-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K diverse SFT data sampled from Tulu3 using ZIP.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-zip-50k.tulu-3-mig-50k
Tulu-3-MIG-50K
Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-mig-50k.tulu-3-random-50k
Tulu-3-Random-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K SFT data randomly sampled from Tulu3.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-random-50k.tulu-3-qdit-50k
Tulu-3-QDIT-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using QDIT.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-qdit-50k.Infinity-Instruct-filteredopenhermes-2.5-pool-annotated
Openhermes-2.5-Pool-Annotated
Annotated Openhermes-2.5. Used as a data pool in MIG. The annotations include #InsTag tags, DEITA scores, and CaR scores.
Dataset Details
Dataset Sources
Repository: https://huggingface.co/datasets/teknium/OpenHermes-2.5
Citation
BibTeX:
@misc{OpenHermes 2.5,
title = {OpenHermes 2.5: An Open Dataset of Synthetic Data for Generalist LLM Assistants},
author = {Teknium},
year = {2023},
publisher… See the full description on the dataset page: https://huggingface.co/datasets/xsample/openhermes-2.5-pool-annotated.CF-Worldeval_bench_xsafety_gpt_4o_guard
Dataset Card for "eval_bench_xsafety_gpt_4o_guard"
More Information needed
tulu-3-car-50k
Tulu-3-CaR-50K
Project | Github | Paper | HuggingFace's collection
This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3 using CaR.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-car-50k.deita-sota-pool-annotated
Deita-Sota-Pool-Annotated
Annotated Deia-Sota-Pool. Used as a data pool in MIG. The annotations include #InsTag tags, DEITA scores, and CaR scores.
Dataset Details
Dataset Sources
Repository: https://huggingface.co/datasets/AndrewZeng/deita_sota_pool
Citation
BibTeX:
deita-sota-mig-6k
Deita-Sota-MIG-6K
Project | Github | Paper | HuggingFace's collection
MIG is an automatic data selection method for instruction tuning.
This dataset includes 50K high-quality and diverse SFT data sampled from Deita-Sota-Pool.
Citation
@article{chen2025mig,
title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space},
author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/deita-sota-mig-6k.tulu-3-ifd-50k
Tulu-3-IFD-50K
Project | Github | Paper | HuggingFace's collection
This dataset includes 50K high-quality SFT data sampled from Tulu3 using IFD.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19
6.71
6.64… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-ifd-50k.tulu-3-instag-50k
Tulu-3-InsTag-50K
Project | Github | Paper | HuggingFace's collection
This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using InsTag.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-instag-50k.eval_bench_xsafety_aegis_gpt_4o
Dataset Card for "eval_bench_xsafety_aegis_gpt_4o"
More Information needed
data_jobs
🧠 data_jobs Dataset
A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse.
Background
I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources.
You can find the full dataset at my app datanerd.tech.
Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/xsarvx/data_jobs.
