CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ToxicityPrompts /XSafetyThe dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. text10K<n<100K0 likes130 downloads1y agoHugging Face02xsample /tulu-3-pool-annotated Tulu-3-Pool-Annotated Project | Github | Paper | HuggingFace's collection Annotated tulu-3-sft-mixture. Used as a data pool in MIG. The annotations include #InsTag tags, DEITA scores, and CaR scores. Dataset Details Tulu3 Dataset Sources Repository: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture Paper [optional]: Tulu 3: Pushing Frontiers in Open Language Model Post-Training MIG Dataset Sources Repository:… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-pool-annotated.text100K<n<1M2 likes55 downloads1y agoHugging Face03xsample /openhermes-2.5-mig-50k Openhermes-2.5-MIG-50K Project | Github | Paper | HuggingFace's collection MIG is an automatic data selection method for instruction tuning. This dataset includes 50K high-quality and diverse SFT data sampled from Openhermes2.5. Citation @article{chen2025mig, title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space}, author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/openhermes-2.5-mig-50k.texttext-generation10K<n<100K1 likes35 downloads1y agoHugging Face04xsample /tulu-3-deita-50k Tulu-3-DEITA-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using DEITA. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-deita-50k.texttext-generation10K<n<100K3 likes35 downloads1y agoHugging Face05xsample /tulu-3-zip-50k Tulu-3-ZIP-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K diverse SFT data sampled from Tulu3 using ZIP. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-zip-50k.texttext-generation10K<n<100K0 likes34 downloads1y agoHugging Face06xsample /tulu-3-mig-50k Tulu-3-MIG-50K Project | Github | Paper | HuggingFace's collection MIG is an automatic data selection method for instruction tuning. This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-mig-50k.texttext-generation10K<n<100K6 likes31 downloads1y agoHugging Face07xsample /tulu-3-random-50k Tulu-3-Random-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K SFT data randomly sampled from Tulu3. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00 59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-random-50k.texttext-generation10K<n<100K0 likes31 downloads1y agoHugging Face08xsample /tulu-3-qdit-50k Tulu-3-QDIT-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using QDIT. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-qdit-50k.texttext-generation10K<n<100K0 likes30 downloads1y agoHugging Face09xsanskarx /Infinity-Instruct-filteredtext100K<n<1M0 likes30 downloads8mo agoHugging Face10xsample /openhermes-2.5-pool-annotated Openhermes-2.5-Pool-Annotated Annotated Openhermes-2.5. Used as a data pool in MIG. The annotations include #InsTag tags, DEITA scores, and CaR scores. Dataset Details Dataset Sources Repository: https://huggingface.co/datasets/teknium/OpenHermes-2.5 Citation BibTeX: @misc{OpenHermes 2.5, title = {OpenHermes 2.5: An Open Dataset of Synthetic Data for Generalist LLM Assistants}, author = {Teknium}, year = {2023}, publisher… See the full description on the dataset page: https://huggingface.co/datasets/xsample/openhermes-2.5-pool-annotated.text1M<n<10M0 likes29 downloads1y agoHugging Face11xsaadsd /CF-Worldtext10K<n<100K0 likes27 downloads5mo agoHugging Face12ToxicityPrompts /eval_bench_xsafety_gpt_4o_guard Dataset Card for "eval_bench_xsafety_gpt_4o_guard" More Information needed text1K<n<10K0 likes24 downloads2y agoHugging Face13xsample /tulu-3-car-50k Tulu-3-CaR-50K Project | Github | Paper | HuggingFace's collection This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3 using CaR. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00 59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-car-50k.texttext-generation10K<n<100K0 likes22 downloads1y agoHugging Face14xsample /deita-sota-pool-annotated Deita-Sota-Pool-Annotated Annotated Deia-Sota-Pool. Used as a data pool in MIG. The annotations include #InsTag tags, DEITA scores, and CaR scores. Dataset Details Dataset Sources Repository: https://huggingface.co/datasets/AndrewZeng/deita_sota_pool Citation BibTeX: text100K<n<1M0 likes21 downloads1y agoHugging Face15xsample /deita-sota-mig-6k Deita-Sota-MIG-6K Project | Github | Paper | HuggingFace's collection MIG is an automatic data selection method for instruction tuning. This dataset includes 50K high-quality and diverse SFT data sampled from Deita-Sota-Pool. Citation @article{chen2025mig, title={MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space}, author={Chen, Yicheng and Li, Yining and Hu, Kai and Ma, Zerun and Ye, Haochen and Chen, Kai}… See the full description on the dataset page: https://huggingface.co/datasets/xsample/deita-sota-mig-6k.texttext-generation1K<n<10K0 likes21 downloads1y agoHugging Face16xsample /tulu-3-ifd-50k Tulu-3-IFD-50K Project | Github | Paper | HuggingFace's collection This dataset includes 50K high-quality SFT data sampled from Tulu3 using IFD. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00 59.19 6.71 6.64… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-ifd-50k.texttext-generation10K<n<100K0 likes17 downloads1y agoHugging Face17xsample /tulu-3-instag-50k Tulu-3-InsTag-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K high-quality and diverse SFT data sampled from Tulu3 using InsTag. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-instag-50k.texttext-generation10K<n<100K0 likes17 downloads1y agoHugging Face18ToxicityPrompts /eval_bench_xsafety_aegis_gpt_4o Dataset Card for "eval_bench_xsafety_aegis_gpt_4o" More Information needed text1K<n<10K0 likes11 downloads2y agoHugging Face19xsarvx /data_jobs 🧠 data_jobs Dataset A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse. Background I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources. You can find the full dataset at my app datanerd.tech. Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/xsarvx/data_jobs.tabular100K<n<1M0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.