datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.DST_Multiwoz21_instruction_Tuning
Dataset Card for "DST_Multiwoz21_instruction_tuning"
More Information needed
TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.brain-instruction-tuningInstruction-tuning_DatasetsTrendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format)
A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles.
Dataset Description
This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.sib200_instructioninvestopedia-instruction-tuning-dataset
Dataset Card for investopedia-instruction-tuning dataset
We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data
and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that
ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-instruction-tuning-dataset.flores_101_instructioncartoonization
Instruction-prompted cartoonization dataset
This dataset was created from 5000 images randomly sampled from the Imagenette dataset. For more
details on how the dataset was created, check out this directory.
Following figure depicts the data preparation workflow:
Known limitations and biases
The dataset was derived from Imagenette, which, in turn, was derived from ImageNet. So, naturally, this
dataset inherits the limitations and biases of ImageNet.… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/cartoonization.Agricultural_pests_and_diseases_instruction_tuning_datalow-level-image-proc
Instruction-prompted low-level image processing dataset
To construct this dataset, we took different number of samples from the following datasets for each task and constructed
a single dataset with prompts added like so:
Task
Prompt
Dataset
Number of samples
Deblurring
“deblur the blurry image”
REDS (train_blur and train_sharp)
1200
Deraining
“derain the image”
Rain13k
686
Denoising
“denoise the noisy image”
SIDD
8
Low-light image enhancement
"enhance the… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/low-level-image-proc.symbolic-instruction-tuning
Symbolic Instruction Tuning
This is the offical repo to host the datasets used in the paper From Zero to Hero: Examining the Power of Symbolic Tasks in Instruction Tuning. The training code can be found in here.
multilingual_instruction_tuningChEMBL_Drug_Instruction_Tuning
Dataset Card for ChEMBL Drug Instruction Tuning
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/alxfgh/ChEMBL_Drug_Instruction_Tuning.v1.1_context_instruction_tuning
Dataset Card for "v1.1_context_instruction_tuning"
More Information needed
multilingual_instruction_tuning_lima_bactrianTCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.Trendyol-Cybersecurity-Instruction-Tuning-Dataset-ConvertedPubChem_Drug_Instruction_Tuninguplimit-instruction-tuning-dataset
Dataset Card for uplimit-instruction-tuning-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset.PubMedVision_InstructionTuning_VQA_splittedforum-instruction-tuning-dataset
Looksmaxxing Forum Dataset
A curated instruction-tuning dataset derived from a large looksmaxxing and aesthetic self-improvement forum,
containing high-density community knowledge on skincare, nutrition, supplementation, and appearance optimization.
Dataset Summary
This dataset was produced by scraping, parsing, cleaning, and LLM-filtering over 1.28 million raw forum posts
down to 77,417 high-quality instruction-response pairs using a multi-stage pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/bediss/forum-instruction-tuning-dataset.multilingual_instruction_tuning_plusmaithili-instruction-tuningswedish-sentiment-instruction-fine-tuning
Dataset Card for "swedish-sentiment-instruction-fine-tuning"
More Information needed
belebele_instructionmedical_instruction_tuninguplimit-instruction-tuning-dataset
Dataset Card for uplimit-instruction-tuning-dataset
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
ipykernel_launcher.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset/raw/main/ipykernel_launcher.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce… See the full description on the dataset page: https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset.marathi-instruction-tuning-alpaca
