datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Instruction-tuning_Datasetscartoonization
Instruction-prompted cartoonization dataset
This dataset was created from 5000 images randomly sampled from the Imagenette dataset. For more
details on how the dataset was created, check out this directory.
Following figure depicts the data preparation workflow:
Known limitations and biases
The dataset was derived from Imagenette, which, in turn, was derived from ImageNet. So, naturally, this
dataset inherits the limitations and biases of ImageNet.… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/cartoonization.low-level-image-proc
Instruction-prompted low-level image processing dataset
To construct this dataset, we took different number of samples from the following datasets for each task and constructed
a single dataset with prompts added like so:
Task
Prompt
Dataset
Number of samples
Deblurring
“deblur the blurry image”
REDS (train_blur and train_sharp)
1200
Deraining
“derain the image”
Rain13k
686
Denoising
“denoise the noisy image”
SIDD
8
Low-light image enhancement
"enhance the… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/low-level-image-proc.PubMedVision_InstructionTuning_VQA_splittedInstruction-Tuning-with-GPT-4-RedPajama-Chat
Instruction Tuning with GPT 4 RedPajama-Chat
This dataset has been converted from the Instruction-Tuning-with-GPT-4 dataset for the purpose of fine-tuning the RedPajama-INCITE-Chat-3B-v1 model.
About Instruction-Tuning-with-GPT-4
English Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
Usage and License Notices
The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Instruction-Tuning-with-GPT-4-RedPajama-Chat.instruction-tuning_empexc-3esFormatted-GSM8K-FewShot-InstructionTuning-OneExample-Gemma2instruction-tuning_EMH-emotional-reactionsFormatted-GSM8K-FewShot-InstructionTuning-OneExampleinstruction-tuning_EMH-interpretationsinstruction-tuning_EMH-explorationsinstruction_tuning_datasets
🧠 Persian Cultural Alignment Dataset for LLMs
This repository contains a high-quality, instruction-following dataset for cultural alignment of large language models (LLMs) in the Persian language. The dataset is curated using hybrid strategies that incorporate culturally grounded generation, multi-turn dialogues, translation, and augmentation methods, making it suitable for SFT, DPO, RLHF, and alignment evaluation.
📚 Dataset Overview
Domain
Methods Used… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/instruction_tuning_datasets.instruction-tuning-translateinstruction_tuning_data
🚀 Load Dataset
from datasets import load_dataset
dataset = load_dataset("shuyuej/instruction_tuning_data", split='train')
print(dataset)
📚 Dataset Description
Data
Size
Link
ChatDoctor
100K
https://www.yunxiangli.top/ChatDoctor/
MedQA
10.2K
https://huggingface.co/datasets/GBaker/MedQA-USMLE-4-options
MedMCQA
183K
https://huggingface.co/datasets/medmcqa
PubmedQA
211K
https://huggingface.co/datasets/pubmed_qa
LiveQA
635… See the full description on the dataset page: https://huggingface.co/datasets/shuyuej/instruction_tuning_data.instruction-tuningFormatted-GSM8K-FewShot-InstructionTuning-TwoExampleInstruction_Tuning_TestingFormatted-GSM8K-FewShot-InstructionTuninginstruction-tuning_ecot_EMH-emotional-reactionsinstruction_tuning_data
