CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01random-long-int /Java_method2test_chatml Java Method to Test ChatML This dataset is based on the methods2test dataset from Microsoft. It follows the ChatML template format: [{'role': '', 'content': ''}, {...}]. Originally, methods2test contains only Java methods at different levels of granularity along with their corresponding test cases. The different focal method segmentations are illustrated here: To simulate a conversation between a Java developer and an AI assistant, I introduce two key parameters: The prompt… See the full description on the dataset page: https://huggingface.co/datasets/random-long-int/Java_method2test_chatml.textquestion-answering100K<n<1M2 likes120 downloads2y agoHugging Face02anshy /Superior-Reasoning-SFT-gpt-oss-120b-random-shuffled Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that… See the full description on the dataset page: https://huggingface.co/datasets/anshy/Superior-Reasoning-SFT-gpt-oss-120b-random-shuffled.texttext-generation100K<n<1M0 likes84 downloads4mo agoHugging Face03benni-ben /random-sentence Random Sentences Dataset This is a random sentence dataset, which was used in SmolBabble-360m, an AI model that spits out random sentences regardless of what is said to it. The dataset has prompt and assistant pairs, and the dataset has 2597 rows. More model info can be found on the model page. Why This was made just for fun, but if you do find an actual use case for this dataset, please open an issue and tell me! texttext-generation1K<n<10K0 likes39 downloads9d agoHugging Face04xsample /tulu-3-random-50k Tulu-3-Random-50K Project | Github | Paper | HuggingFace's collection This dataset is a baseline of MIG. It includes 50K SFT data randomly sampled from Tulu3. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00 59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-random-50k.texttext-generation10K<n<100K0 likes28 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.