CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigcode /self-oss-instruct-sc2-exec-filter-50kFinal self-alignment training dataset for StarCoder2-Instruct. seed: Contains the seed Python function concepts: Contains the concepts generated from the seed instruction: Contains the instruction generated from the concepts response: Contains the execution-validated response to the instruction This dataset utilizes seed Python functions derived from the MultiPL-T pipeline. text10K<n<100K108 likes39k downloads2y agoHugging Face02HuggingFaceH4 /testing_self_instruct_small Dataset Card for "testing_self_instruct_small" More Information needed textn<1K2 likes16k downloads3y agoHugging Face03MemGPT /MSC-Self-Instruct MemGPT This is the self-instruct dataset of MSC conversations used for MemGPT paper. For more information please refer to memgpt.ai The MSC dataset is a multi-round human conversations. In this dataset, our goal is to come up with a conversation opener, that is personalized to the user by referencing topics from the previous conversations. These were generated while evaluating MemGPT. textn<1K13 likes1.1k downloads3y agoHugging Face04codeparrot /self-instruct-starcoder Self-instruct-starcoder Summary Self-instruct-starcoder is a dataset that was generated by prompting starcoder to generate new instructions based on some human-written seed instructions. The underlying process is explained in the paper self-instruct. This algorithm gave birth to famous machine generated datasets such as Alpaca and Code Alpaca which are two datasets obtained by prompting OpenAI text-davinci-003 engine. Our approach While our method is… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/self-instruct-starcoder.text1K<n<10K64 likes690 downloads3y agoHugging Face05yizhongw /self_instructSelf-Instruct is a dataset that contains 52k instructions, paired with 82K instance inputs and outputs. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.text100K<n<1M194 likes586 downloads4y agoHugging Face06zwq2018 /Multi-modal-Self-instruct Dataset Description Paper Information Dataset Examples Leaderboard Dataset Usage Data Downloading Data Format Evaluation Citation You can download the zip dataset directly, and both train and test subsets are collected in Multi-modal-Self-instruct.zip. Dataset Description Multi-Modal Self-Instruct dataset utilizes large language models and their code capabilities to synthesize massive abstract images and visual reasoning instructions across daily scenarios. This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct.imagemultiple-choice10K<n<100K34 likes481 downloads2y agoHugging Face07HuggingFaceTB /self-oss-instruct-sc2-H4StarCoder2-Self-Instruct-OSS-50k dataset formatted to be compatible with the alignement-handbook for SFT. text10K<n<100K6 likes218 downloads2y agoHugging Face08laion /dev_set_v2_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260825_134035text1K<n<10K0 likes215 downloads27d agoHugging Face09laion /terminal_bench_2_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260f49e33datext1K<n<10K0 likes213 downloads23d agoHugging Face10laion /swebench_verified_random_100_folders_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_c2a08420text1K<n<10K0 likes211 downloads23d agoHugging Face11open-athena /selfinstruct-naive-sandboxes-2-verified-qwen3.5-122b-131k-opencode-tracestext1K<n<10K0 likes193 downloads2mo agoHugging Face12open-athena /rl__24GPU_shaped__selfinstruct-naive-sandboxes-2-verified__exp_tas_optimal_comb__40-0text10K<n<100K0 likes175 downloads6mo agoHugging Face13HuggingFaceH4 /self_instructThis dataset splits the original Self-instruct dataset into training (90%) and test (10%). texttext-generation10K<n<100K10 likes158 downloads3y agoHugging Face14open-athena /a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verifiedtext10K<n<100K0 likes157 downloads4mo agoHugging Face15DCAgent3 /bfcl_parity_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192648textn<1K0 likes153 downloads4mo agoHugging Face16DCAgent3 /swebench_verified_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_2d90a6650text1K<n<10K0 likes151 downloads4mo agoHugging Face17DCAgent3 /medagentbench_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_2026061f18c01text1K<n<10K0 likes137 downloads4mo agoHugging Face18DCAgent3 /terminal_bench_2_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_204c8a44f2textn<1K0 likes135 downloads4mo agoHugging Face19jiachenli-ucsb /self-oss-instruct-sc2-exec-filter-prompt-codes-test-50ktext10K<n<100K0 likes130 downloads2y agoHugging Face20HINT-lab /Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration. 🔥 Updates [2025-3-3]: We released our paper. [2025-2-25]: We released our codes, models and datasets. 🏴󠁶󠁵󠁭󠁡󠁰󠁿 Overview We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.tabularquestion-answering100K<n<1M0 likes130 downloads2y agoHugging Face21DCAgent3 /gaia_127_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192659textn<1K0 likes130 downloads4mo agoHugging Face22DCAgent3 /swebench_verified_random_100_folders_a3_rl_DCAgent_selfinstruct_naive_sandboxes6b437cbatextn<1K0 likes130 downloads4mo agoHugging Face23DCAgent2 /dev_set_v2_rl__24GPU_shaped__selfinstruct_naive_sandboxes_2_verified__exp_tas_o57316c9atextn<1K0 likes127 downloads6mo agoHugging Face24DCAgent3 /dev_set_v2_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192604textn<1K0 likes126 downloads4mo agoHugging Face25GenAIDevTOProd /multimodal-vqa-self-instruct-enriched Multimodal VQA – Self-Instruct-enriched Overview This dataset is an enriched, cleaned, and metadata-enhanced version of zwq2018/Multi-modal-Self-instruct.It pairs images with natural language questions and answers, making it ideal for Vision-Language Model (VLM) training, benchmarking, and instruction tuning. Dataset Summary Total samples: 75,000+ (64,796 train, 11,193 test) Modalities: Image + Text (Questions) + Text (Answers) Task Types: Visual Question… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/multimodal-vqa-self-instruct-enriched.image10K<n<100K1 likes122 downloads1y agoHugging Face26DCAgent3 /aider_polyglot_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_202690541bcetextn<1K0 likes122 downloads4mo agoHugging Face27sambanovasystems /x-self-instruct-seed-32 Dataset Card for xOA22 - Multilingual Prompts from OpenAssistant Dataset Summary x-self-instruct-seed-32 consists of 32 prompts chosen out of the 252 prompts in the self-instruct-seed dataset from the Self-Instruct paper. These 32 prompts were filtered out according to the following criteria: Should be natural in a chat setting Therefore, we filter out any prompts with "few-shot examples", as these are all instruction prompts that we consider unnatural in a chat setting… See the full description on the dataset page: https://huggingface.co/datasets/sambanovasystems/x-self-instruct-seed-32.textn<1K1 likes121 downloads3y agoHugging Face28DCAgent2 /terminal_bench_2_rl__24GPU_shaped__selfinstruct_naive_sandboxes_2_verified__exp93c24543textn<1K0 likes121 downloads6mo agoHugging Face29DCAgent3 /financeagent_terminal_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_5349063btextn<1K0 likes121 downloads4mo agoHugging Face30fwnlp /self-instruct-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models 🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct) Disclaimer The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk. Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/self-instruct-safety-alignment.text10K<n<100K4 likes115 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.