CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K134 likes5.7k downloads1y agoHugging Face02ShuaiYang03 /VLA_Instruction_TuningThis repository contains the VLA-IT dataset, a curated 650K-sample Vision-Language-Action Instruction Tuning dataset, and the SimplerEnv-Instruct benchmark. These are presented in the paper InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation. The dataset is designed to enable robots to integrate multimodal reasoning with precise action generation, preserving the flexible reasoning of large vision-language models while delivering leading manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/VLA_Instruction_Tuning.robotics4 likes5.3k downloads1y agoHugging Face03AtheerAlgherairy /DST_Multiwoz21_instruction_Tuning Dataset Card for "DST_Multiwoz21_instruction_tuning" More Information needed text10K<n<100K0 likes761 downloads3y agoHugging Face04iamplus /Instruction_TuningFiles Contents Details : Post-Process Code Info : data_process.py iamai_seed_tasks_v1.csv : IAMAI's seed tasks - Version 1 (879) Total Dataset Size : 879 =============================================================================================== iamai_v1.csv : Instruction Tuning Dataset collected using seeds from iamai_seed_tasks_v1.csv and ChatGPT API for both prompts and outputs (~248k) Total Dataset Size : ~248k iamai_summarization_v1.csv : Article Summarization dataset (both… See the full description on the dataset page: https://huggingface.co/datasets/iamplus/Instruction_Tuning.2 likes453 downloads3y agoHugging Face05appledora /DORI-instruction-tuning-dataset DORI Spatial Reasoning Instruction Dataset Dataset Description This dataset contains instruction tuning data for spatial reasoning tasks across multiple question types and visual datasets. Dataset Structure Dataset Splits train: 26,626 samples test: 6,672 samples Total: 33,298 samples Question Types q1 q2 q3 q4 q5 q6 q7 Source Datasets 3d_future cityscapes coco coco_space_sea get_3d jta kitti nocs_real objectron… See the full description on the dataset page: https://huggingface.co/datasets/appledora/DORI-instruction-tuning-dataset.imagevisual-question-answering10K<n<100K1 likes451 downloads1y agoHugging Face06ChiyuSONG /dynamics-of-instruction-tuning 💻 [Github Repo] • 📃 [Paper] • 👀 [Preview] Update 12/01/23: Corrected ambiguous choices in the validation and test sets of the role-play chat data. Overview We introduce DoIT, a collection of over 40k human-curated instruction-output pairs in Chinese. This dataset is organized into ten representative ability categories: (1) STEM subject - Biology, (2) Humanity subject - History, (3) Code Generation, (4) Creative Writing, (5) Language proficiency - Chinese, (6)… See the full description on the dataset page: https://huggingface.co/datasets/ChiyuSONG/dynamics-of-instruction-tuning.text-generation5 likes433 downloads2y agoHugging Face07FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes428 downloads1y agoHugging Face08BoltzmachineQ /brain-instruction-tuningtext100K<n<1M3 likes427 downloads1y agoHugging Face09ldbb123 /Instruction-tuning_Datasetstext1M<n<10M0 likes262 downloads2y agoHugging Face10tuandunghcmut /Trendyol-Cybersecurity-Instruction-Tuning-Datasetgated Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format) A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles. Dataset Description This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K2 likes216 downloads1y agoHugging Face11mbzuai-ugrip-statement-tuning /sib200_instructiontext10K<n<100K0 likes208 downloads2y agoHugging Face12FinLang /investopedia-instruction-tuning-dataset Dataset Card for investopedia-instruction-tuning dataset We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-instruction-tuning-dataset.text100K<n<1M23 likes201 downloads2y agoHugging Face13mbzuai-ugrip-statement-tuning /flores_101_instructiontext100K<n<1M0 likes183 downloads2y agoHugging Face14instruction-tuning-sd /cartoonization Instruction-prompted cartoonization dataset This dataset was created from 5000 images randomly sampled from the Imagenette dataset. For more details on how the dataset was created, check out this directory. Following figure depicts the data preparation workflow: Known limitations and biases The dataset was derived from Imagenette, which, in turn, was derived from ImageNet. So, naturally, this dataset inherits the limitations and biases of ImageNet.… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/cartoonization.imageimage-to-image1K<n<10K21 likes168 downloads3y agoHugging Face15Agri-LLaVA-Anonymous /Agricultural_pests_and_diseases_instruction_tuning_datatext1K<n<10K2 likes158 downloads2y agoHugging Face16SmartQHSE /hse-instruction-tuning SmartQHSE HSE Instruction-Tuning Corpus Alpaca-style instruction-tuning variant of the SmartQHSE HSE Q&A Corpus. Each row has the canonical fine-tuning schema: { "instruction": "What is OSHA Process Safety Management 1910.119?", "input": "", "output": "<authoritative long-form answer with citations>", "category": "us-osha", "source_url": "https://www.smartqhse.com/answers/<slug>" } Suitable for LoRA / SFT training of HSE-domain LLMs and RAG systems. Citation (preferred… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-instruction-tuning.text-generationn<1K0 likes150 downloads5mo agoHugging Face17instruction-tuning-sd /low-level-image-proc Instruction-prompted low-level image processing dataset To construct this dataset, we took different number of samples from the following datasets for each task and constructed a single dataset with prompts added like so: Task Prompt Dataset Number of samples Deblurring “deblur the blurry image” REDS (train_blur and train_sharp) 1200 Deraining “derain the image” Rain13k 686 Denoising “denoise the noisy image” SIDD 8 Low-light image enhancement "enhance the… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/low-level-image-proc.imageimage-to-image1K<n<10K9 likes141 downloads3y agoHugging Face18sail /symbolic-instruction-tuning Symbolic Instruction Tuning This is the offical repo to host the datasets used in the paper From Zero to Hero: Examining the Power of Symbolic Tasks in Instruction Tuning. The training code can be found in here. text100K<n<1M16 likes134 downloads3y agoHugging Face19kuyesu22 /multilingual_instruction_tuningtext100K<n<1M0 likes132 downloads1y agoHugging Face20dougalldeepmind /2026-08-04-table2-instruction-tuning-9284-filtered-8192 Table 2 instruction-tuning mixture — spec-filtered, 8192-safe (9,284 examples) The paper's Table 2 instruction-tuning mixture, spec-filtered, with the single row that cannot fit an 8,192-token window removed. No difficult-advice data — this is the general instruction-tuning half on its own. field value experiment Table 2 instruction-tuning mixture for the Teaching Claude Why replication, filtered for spec misalignment and trimmed to fit max_seq_len 8192… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-table2-instruction-tuning-9284-filtered-8192.text-generation0 likes132 downloads1mo agoHugging Face21SciReason /SciLM-Instruction_Tuning SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines This repo contains the instruction-tuning data of SciReasoner. 3 likes123 downloads1y agoHugging Face22alxfgh /ChEMBL_Drug_Instruction_Tuning Dataset Card for ChEMBL Drug Instruction Tuning Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/alxfgh/ChEMBL_Drug_Instruction_Tuning.textquestion-answering100K<n<1M15 likes122 downloads3y agoHugging Face23tyzhu /v1.1_context_instruction_tuning Dataset Card for "v1.1_context_instruction_tuning" More Information needed text100K<n<1M1 likes116 downloads3y agoHugging Face24junkim100 /multilingual_instruction_tuning_lima_bactriantext100K<n<1M0 likes109 downloads1y agoHugging Face25CarsonnnNN /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M0 likes109 downloads7mo agoHugging Face26ChavyvAkvar /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-Convertedtext10K<n<100K1 likes108 downloads1y agoHugging Face27dougalldeepmind /2026-08-04-table2-instruction-tuning-mixture-spec-filtered Table 2 instruction-tuning mixture, spec-filtered A reproduction of the paper's Table 2 instruction-tuning mixture at its exact per-source sample counts, plus an LLM spec-alignment filter and the per-sample judge verdicts, so the filter can be re-cut at any threshold without paying to re-judge. field value experiment Table 2 instruction-tuning mixture for the Teaching Claude Why replication, filtered for spec misalignment date_generated 2026-08-04 constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-table2-instruction-tuning-mixture-spec-filtered.text-generation0 likes102 downloads1mo agoHugging Face28alxfgh /PubChem_Drug_Instruction_Tuningtext10K<n<100K11 likes98 downloads3y agoHugging Face29landedmover /uplimit-instruction-tuning-dataset Dataset Card for uplimit-instruction-tuning-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset.textn<1K0 likes89 downloads2y agoHugging Face30ChuGyouk /PubMedVision_InstructionTuning_VQA_splittedtext100K<n<1M1 likes87 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.