CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xsample /openhermes-2.5-mig-50k Openhermes-2.5-MIG-50K Project | Github | Paper | HuggingFace's collection MIG is an automatic data selection method for instruction tuning. This dataset includes 50K high-quality and diverse SFT data sampled from Openhermes2.5. Citation @article{chen2025mig, title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space}, author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/openhermes-2.5-mig-50k.texttext-generation10K<n<100K1 likes35 downloads1y agoHugging Face02xsample /tulu-3-deita-50k Tulu-3-DEITA-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using DEITA. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-deita-50k.texttext-generation10K<n<100K3 likes35 downloads1y agoHugging Face03xsample /tulu-3-zip-50k Tulu-3-ZIP-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K diverse SFT data sampled from Tulu3 using ZIP. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-zip-50k.texttext-generation10K<n<100K0 likes34 downloads1y agoHugging Face04xsample /tulu-3-mig-50k Tulu-3-MIG-50K Project | Github | Paper | HuggingFace's collection MIG is an automatic data selection method for instruction tuning. This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-mig-50k.texttext-generation10K<n<100K6 likes31 downloads1y agoHugging Face05xsample /tulu-3-random-50k Tulu-3-Random-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K SFT data randomly sampled from Tulu3. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00 59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-random-50k.texttext-generation10K<n<100K0 likes31 downloads1y agoHugging Face06xsample /tulu-3-qdit-50k Tulu-3-QDIT-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using QDIT. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-qdit-50k.texttext-generation10K<n<100K0 likes30 downloads1y agoHugging Face07xsample /tulu-3-car-50k Tulu-3-CaR-50K Project | Github | Paper | HuggingFace's collection This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3 using CaR. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00 59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-car-50k.texttext-generation10K<n<100K0 likes22 downloads1y agoHugging Face08xsample /deita-sota-mig-6k Deita-Sota-MIG-6K Project | Github | Paper | HuggingFace's collection MIG is an automatic data selection method for instruction tuning. This dataset includes 50K high-quality and diverse SFT data sampled from Deita-Sota-Pool. Citation @article{chen2025mig, title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space}, author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/deita-sota-mig-6k.texttext-generation1K<n<10K0 likes21 downloads1y agoHugging Face09xsample /tulu-3-ifd-50k Tulu-3-IFD-50K Project | Github | Paper | HuggingFace's collection This dataset includes 50K high-quality SFT data sampled from Tulu3 using IFD. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00 59.19 6.71 6.64… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-ifd-50k.texttext-generation10K<n<100K0 likes17 downloads1y agoHugging Face10xsample /tulu-3-instag-50k Tulu-3-InsTag-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using InsTag. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-instag-50k.texttext-generation10K<n<100K0 likes17 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.