datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DST_Multiwoz21_instruction_Tuning
Dataset Card for "DST_Multiwoz21_instruction_tuning"
More Information needed
Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format)
A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles.
Dataset Description
This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.sib200_instructionflores_101_instructioncartoonization
Instruction-prompted cartoonization dataset
This dataset was created from 5000 images randomly sampled from the Imagenette dataset. For more
details on how the dataset was created, check out this directory.
Following figure depicts the data preparation workflow:
Known limitations and biases
The dataset was derived from Imagenette, which, in turn, was derived from ImageNet. So, naturally, this
dataset inherits the limitations and biases of ImageNet.… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/cartoonization.low-level-image-proc
Instruction-prompted low-level image processing dataset
To construct this dataset, we took different number of samples from the following datasets for each task and constructed
a single dataset with prompts added like so:
Task
Prompt
Dataset
Number of samples
Deblurring
“deblur the blurry image”
REDS (train_blur and train_sharp)
1200
Deraining
“derain the image”
Rain13k
686
Denoising
“denoise the noisy image”
SIDD
8
Low-light image enhancement
"enhance the… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/low-level-image-proc.multilingual_instruction_tuningv1.1_context_instruction_tuning
Dataset Card for "v1.1_context_instruction_tuning"
More Information needed
multilingual_instruction_tuning_lima_bactrianTrendyol-Cybersecurity-Instruction-Tuning-Dataset-Converteduplimit-instruction-tuning-dataset
Dataset Card for uplimit-instruction-tuning-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset.PubMedVision_InstructionTuning_VQA_splittedmultilingual_instruction_tuning_plusswedish-sentiment-instruction-fine-tuning
Dataset Card for "swedish-sentiment-instruction-fine-tuning"
More Information needed
belebele_instructionuplimit-instruction-tuning-dataset
Dataset Card for uplimit-instruction-tuning-dataset
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
ipykernel_launcher.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset/raw/main/ipykernel_launcher.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce… See the full description on the dataset page: https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset.marathi-instruction-tuning-alpacawiki_lingua_instructionpersian-instruction-tuning-jsonlxquad_instructionxcodah_instructionpersian-instruction-tuningmultilingual_instruction_tuningsentiments_instructionluganda_instruction_tuning
Paper and Citation
Paper: Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages
@misc{toukmaji2025prompttranslatefinetunereinitialize,
title={Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages},
author={Christopher Toukmaji and Jeffrey Flanigan},
year={2025},
eprint={2506.19187}… See the full description on the dataset page: https://huggingface.co/datasets/ChrisToukmaji/luganda_instruction_tuning.v1.1_id0.2_context_instruction_tuning
Dataset Card for "v1.1_id0.2_context_instruction_tuning"
More Information needed
Tumbuka_Instruction_Tuningxcsqa_instructioncs_instruction_tuning_collection
Dataset Card for Czech Instruction Tuning Collection
This dataset is a collection for instruction tuning of LLMs in Czech language.
Dataset Details
Dataset Description
Curated by: Artificial Intelligence Center, FEE, CTU in Prague
Language(s) (NLP): Czech (cs, ces)
License: cc-by-nc-4.0
Dataset Sources
The data points in the dataset were collected from following sources:
MURI-IT - supernatural instructions, WikiHow, Reverse instructions… See the full description on the dataset page: https://huggingface.co/datasets/ctu-aic/cs_instruction_tuning_collection.xlwic_instruction
