CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K134 likes5.7k downloads1y agoHugging Face02AtheerAlgherairy /DST_Multiwoz21_instruction_Tuning Dataset Card for "DST_Multiwoz21_instruction_tuning" More Information needed text10K<n<100K0 likes761 downloads3y agoHugging Face03FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes428 downloads1y agoHugging Face04BoltzmachineQ /brain-instruction-tuningtext100K<n<1M3 likes427 downloads1y agoHugging Face05ldbb123 /Instruction-tuning_Datasetstext1M<n<10M0 likes262 downloads2y agoHugging Face06tuandunghcmut /Trendyol-Cybersecurity-Instruction-Tuning-Datasetgated Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format) A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles. Dataset Description This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K2 likes216 downloads1y agoHugging Face07mbzuai-ugrip-statement-tuning /sib200_instructiontext10K<n<100K0 likes208 downloads2y agoHugging Face08FinLang /investopedia-instruction-tuning-dataset Dataset Card for investopedia-instruction-tuning dataset We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-instruction-tuning-dataset.text100K<n<1M23 likes201 downloads2y agoHugging Face09mbzuai-ugrip-statement-tuning /flores_101_instructiontext100K<n<1M0 likes183 downloads2y agoHugging Face10instruction-tuning-sd /cartoonization Instruction-prompted cartoonization dataset This dataset was created from 5000 images randomly sampled from the Imagenette dataset. For more details on how the dataset was created, check out this directory. Following figure depicts the data preparation workflow: Known limitations and biases The dataset was derived from Imagenette, which, in turn, was derived from ImageNet. So, naturally, this dataset inherits the limitations and biases of ImageNet.… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/cartoonization.imageimage-to-image1K<n<10K21 likes168 downloads3y agoHugging Face11Agri-LLaVA-Anonymous /Agricultural_pests_and_diseases_instruction_tuning_datatext1K<n<10K2 likes158 downloads2y agoHugging Face12instruction-tuning-sd /low-level-image-proc Instruction-prompted low-level image processing dataset To construct this dataset, we took different number of samples from the following datasets for each task and constructed a single dataset with prompts added like so: Task Prompt Dataset Number of samples Deblurring “deblur the blurry image” REDS (train_blur and train_sharp) 1200 Deraining “derain the image” Rain13k 686 Denoising “denoise the noisy image” SIDD 8 Low-light image enhancement "enhance the… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/low-level-image-proc.imageimage-to-image1K<n<10K9 likes141 downloads3y agoHugging Face13sail /symbolic-instruction-tuning Symbolic Instruction Tuning This is the offical repo to host the datasets used in the paper From Zero to Hero: Examining the Power of Symbolic Tasks in Instruction Tuning. The training code can be found in here. text100K<n<1M16 likes134 downloads3y agoHugging Face14kuyesu22 /multilingual_instruction_tuningtext100K<n<1M0 likes132 downloads1y agoHugging Face15alxfgh /ChEMBL_Drug_Instruction_Tuning Dataset Card for ChEMBL Drug Instruction Tuning Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/alxfgh/ChEMBL_Drug_Instruction_Tuning.textquestion-answering100K<n<1M15 likes122 downloads3y agoHugging Face16tyzhu /v1.1_context_instruction_tuning Dataset Card for "v1.1_context_instruction_tuning" More Information needed text100K<n<1M1 likes116 downloads3y agoHugging Face17junkim100 /multilingual_instruction_tuning_lima_bactriantext100K<n<1M0 likes109 downloads1y agoHugging Face18CarsonnnNN /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M0 likes109 downloads7mo agoHugging Face19ChavyvAkvar /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-Convertedtext10K<n<100K1 likes108 downloads1y agoHugging Face20alxfgh /PubChem_Drug_Instruction_Tuningtext10K<n<100K11 likes98 downloads3y agoHugging Face21landedmover /uplimit-instruction-tuning-dataset Dataset Card for uplimit-instruction-tuning-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset.textn<1K0 likes89 downloads2y agoHugging Face22ChuGyouk /PubMedVision_InstructionTuning_VQA_splittedtext100K<n<1M1 likes87 downloads2y agoHugging Face23bediss /forum-instruction-tuning-dataset Looksmaxxing Forum Dataset A curated instruction-tuning dataset derived from a large looksmaxxing and aesthetic self-improvement forum, containing high-density community knowledge on skincare, nutrition, supplementation, and appearance optimization. Dataset Summary This dataset was produced by scraping, parsing, cleaning, and LLM-filtering over 1.28 million raw forum posts down to 77,417 high-quality instruction-response pairs using a multi-stage pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/bediss/forum-instruction-tuning-dataset.texttext-generation10K<n<100K1 likes79 downloads3mo agoHugging Face24treasure4l /multilingual_instruction_tuning_plustext100K<n<1M0 likes79 downloads1mo agoHugging Face25Bansal123 /maithili-instruction-tuningtext1K<n<10K2 likes74 downloads7mo agoHugging Face26filopedraz /swedish-sentiment-instruction-fine-tuning Dataset Card for "swedish-sentiment-instruction-fine-tuning" More Information needed text100K<n<1M2 likes70 downloads3y agoHugging Face27mbzuai-ugrip-statement-tuning /belebele_instructiontext10K<n<100K0 likes68 downloads2y agoHugging Face28Star-gazer /medical_instruction_tuningtextquestion-answering10K<n<100K0 likes65 downloads2y agoHugging Face29BelarminoF /uplimit-instruction-tuning-dataset Dataset Card for uplimit-instruction-tuning-dataset This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: ipykernel_launcher.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset/raw/main/ipykernel_launcher.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce… See the full description on the dataset page: https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset.textn<1K0 likes63 downloads2y agoHugging Face30smallstepai /marathi-instruction-tuning-alpacatext10K<n<100K2 likes62 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.