CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AtheerAlgherairy /DST_Multiwoz21_instruction_Tuning Dataset Card for "DST_Multiwoz21_instruction_tuning" More Information needed text10K<n<100K0 likes761 downloads3y agoHugging Face02tuandunghcmut /Trendyol-Cybersecurity-Instruction-Tuning-Datasetgated Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format) A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles. Dataset Description This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K2 likes216 downloads1y agoHugging Face03mbzuai-ugrip-statement-tuning /sib200_instructiontext10K<n<100K0 likes208 downloads2y agoHugging Face04mbzuai-ugrip-statement-tuning /flores_101_instructiontext100K<n<1M0 likes183 downloads2y agoHugging Face05instruction-tuning-sd /cartoonization Instruction-prompted cartoonization dataset This dataset was created from 5000 images randomly sampled from the Imagenette dataset. For more details on how the dataset was created, check out this directory. Following figure depicts the data preparation workflow: Known limitations and biases The dataset was derived from Imagenette, which, in turn, was derived from ImageNet. So, naturally, this dataset inherits the limitations and biases of ImageNet.… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/cartoonization.imageimage-to-image1K<n<10K21 likes168 downloads3y agoHugging Face06instruction-tuning-sd /low-level-image-proc Instruction-prompted low-level image processing dataset To construct this dataset, we took different number of samples from the following datasets for each task and constructed a single dataset with prompts added like so: Task Prompt Dataset Number of samples Deblurring “deblur the blurry image” REDS (train_blur and train_sharp) 1200 Deraining “derain the image” Rain13k 686 Denoising “denoise the noisy image” SIDD 8 Low-light image enhancement "enhance the… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/low-level-image-proc.imageimage-to-image1K<n<10K9 likes141 downloads3y agoHugging Face07kuyesu22 /multilingual_instruction_tuningtext100K<n<1M0 likes132 downloads1y agoHugging Face08tyzhu /v1.1_context_instruction_tuning Dataset Card for "v1.1_context_instruction_tuning" More Information needed text100K<n<1M1 likes116 downloads3y agoHugging Face09junkim100 /multilingual_instruction_tuning_lima_bactriantext100K<n<1M0 likes109 downloads1y agoHugging Face10ChavyvAkvar /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-Convertedtext10K<n<100K1 likes108 downloads1y agoHugging Face11landedmover /uplimit-instruction-tuning-dataset Dataset Card for uplimit-instruction-tuning-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset.textn<1K0 likes89 downloads2y agoHugging Face12ChuGyouk /PubMedVision_InstructionTuning_VQA_splittedtext100K<n<1M1 likes87 downloads2y agoHugging Face13treasure4l /multilingual_instruction_tuning_plustext100K<n<1M0 likes79 downloads1mo agoHugging Face14filopedraz /swedish-sentiment-instruction-fine-tuning Dataset Card for "swedish-sentiment-instruction-fine-tuning" More Information needed text100K<n<1M2 likes70 downloads3y agoHugging Face15mbzuai-ugrip-statement-tuning /belebele_instructiontext10K<n<100K0 likes68 downloads2y agoHugging Face16BelarminoF /uplimit-instruction-tuning-dataset Dataset Card for uplimit-instruction-tuning-dataset This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: ipykernel_launcher.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset/raw/main/ipykernel_launcher.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce… See the full description on the dataset page: https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset.textn<1K0 likes63 downloads2y agoHugging Face17smallstepai /marathi-instruction-tuning-alpacatext10K<n<100K2 likes62 downloads3y agoHugging Face18mbzuai-ugrip-statement-tuning /wiki_lingua_instructiontext100K<n<1M0 likes61 downloads2y agoHugging Face19PersianML /persian-instruction-tuning-jsonltext100K<n<1M0 likes60 downloads2mo agoHugging Face20mbzuai-ugrip-statement-tuning /xquad_instructiontext10K<n<100K0 likes55 downloads2y agoHugging Face21mbzuai-ugrip-statement-tuning /xcodah_instructiontext1K<n<10K0 likes53 downloads2y agoHugging Face22PersianML /persian-instruction-tuningtext100K<n<1M0 likes51 downloads2mo agoHugging Face23treasure4l /multilingual_instruction_tuningtext100K<n<1M0 likes50 downloads1mo agoHugging Face24mbzuai-ugrip-statement-tuning /sentiments_instructiontext100K<n<1M0 likes43 downloads2y agoHugging Face25ChrisToukmaji /luganda_instruction_tuning Paper and Citation Paper: Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages @misc{toukmaji2025prompttranslatefinetunereinitialize, title={Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages}, author={Christopher Toukmaji and Jeffrey Flanigan}, year={2025}, eprint={2506.19187}… See the full description on the dataset page: https://huggingface.co/datasets/ChrisToukmaji/luganda_instruction_tuning.text1K<n<10K0 likes42 downloads1y agoHugging Face26tyzhu /v1.1_id0.2_context_instruction_tuning Dataset Card for "v1.1_id0.2_context_instruction_tuning" More Information needed text100K<n<1M0 likes40 downloads3y agoHugging Face27Mwanzau /Tumbuka_Instruction_Tuningtext1M<n<10M0 likes40 downloads2mo agoHugging Face28mbzuai-ugrip-statement-tuning /xcsqa_instructiontext10K<n<100K0 likes39 downloads2y agoHugging Face29ctu-aic /cs_instruction_tuning_collection Dataset Card for Czech Instruction Tuning Collection This dataset is a collection for instruction tuning of LLMs in Czech language. Dataset Details Dataset Description Curated by: Artificial Intelligence Center, FEE, CTU in Prague Language(s) (NLP): Czech (cs, ces) License: cc-by-nc-4.0 Dataset Sources The data points in the dataset were collected from following sources: MURI-IT - supernatural instructions, WikiHow, Reverse instructions… See the full description on the dataset page: https://huggingface.co/datasets/ctu-aic/cs_instruction_tuning_collection.texttext-generation100K<n<1M1 likes38 downloads1y agoHugging Face30mbzuai-ugrip-statement-tuning /xlwic_instructiontext10K<n<100K0 likes38 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.