datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.dynamics-of-instruction-tuning
💻 [Github Repo] • 📃 [Paper] • 👀 [Preview]
Update
12/01/23: Corrected ambiguous choices in the validation and test sets of the role-play chat data.
Overview
We introduce DoIT, a collection of over 40k human-curated instruction-output pairs in Chinese. This dataset is organized into ten representative ability categories: (1) STEM subject - Biology, (2) Humanity subject - History, (3) Code Generation, (4) Creative Writing, (5) Language proficiency - Chinese, (6)… See the full description on the dataset page: https://huggingface.co/datasets/ChiyuSONG/dynamics-of-instruction-tuning.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format)
A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles.
Dataset Description
This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.hse-instruction-tuning
SmartQHSE HSE Instruction-Tuning Corpus
Alpaca-style instruction-tuning variant of the SmartQHSE HSE Q&A Corpus.
Each row has the canonical fine-tuning schema:
{
"instruction": "What is OSHA Process Safety Management 1910.119?",
"input": "",
"output": "<authoritative long-form answer with citations>",
"category": "us-osha",
"source_url": "https://www.smartqhse.com/answers/<slug>"
}
Suitable for LoRA / SFT training of HSE-domain LLMs and RAG systems.
Citation (preferred… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-instruction-tuning.2026-08-04-table2-instruction-tuning-9284-filtered-8192
Table 2 instruction-tuning mixture — spec-filtered, 8192-safe (9,284 examples)
The paper's Table 2 instruction-tuning mixture, spec-filtered, with the single row that
cannot fit an 8,192-token window removed. No difficult-advice data — this is the
general instruction-tuning half on its own.
field
value
experiment
Table 2 instruction-tuning mixture for the Teaching Claude Why replication, filtered for spec misalignment and trimmed to fit max_seq_len 8192… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-table2-instruction-tuning-9284-filtered-8192.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.2026-08-04-table2-instruction-tuning-mixture-spec-filtered
Table 2 instruction-tuning mixture, spec-filtered
A reproduction of the paper's Table 2 instruction-tuning mixture at its exact per-source
sample counts, plus an LLM spec-alignment filter and the per-sample judge verdicts, so
the filter can be re-cut at any threshold without paying to re-judge.
field
value
experiment
Table 2 instruction-tuning mixture for the Teaching Claude Why replication, filtered for spec misalignment
date_generated
2026-08-04
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-table2-instruction-tuning-mixture-spec-filtered.forum-instruction-tuning-dataset
Looksmaxxing Forum Dataset
A curated instruction-tuning dataset derived from a large looksmaxxing and aesthetic self-improvement forum,
containing high-density community knowledge on skincare, nutrition, supplementation, and appearance optimization.
Dataset Summary
This dataset was produced by scraping, parsing, cleaning, and LLM-filtering over 1.28 million raw forum posts
down to 77,417 high-quality instruction-response pairs using a multi-stage pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/bediss/forum-instruction-tuning-dataset.OneLLM_InstructionTuning
Data
Data Format
All finetuning data are converted into multi-turn conversation format. The .json file contains a list of training samples, where each sample contains the following keys: id, image and conversations. For example,
{'id': '000000033471', 'image': 'InstructionTuning/image/coco/train2017/000000033471.jpg', 'conversations': [{'from': 'human', 'value': 'What are the colors of the bus in the image?'}, {'from': 'gpt', 'value': 'The bus in the image is white and… See the full description on the dataset page: https://huggingface.co/datasets/csuhan/OneLLM_InstructionTuning.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/Soban1234/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.snorkel-curated-instruction-tuningPlease check out our Blog Post - How we built a better GenAI with programmatic data development for more details!
Summary
snorkel-curated-instruction-tuning is a curated dataset that consists of high-quality instruction-response pairs.
These pairs were programmatically filtered with weak supervision from open-source datasets Databricks Dolly-15k,
Open Assistant,
and Helpful Instructions.
To enhance the dataset, we also programmatically classified each instruction based on the… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/snorkel-curated-instruction-tuning.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ahmadkaab/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/andycoco1128/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.cs_instruction_tuning_collection
Dataset Card for Czech Instruction Tuning Collection
This dataset is a collection for instruction tuning of LLMs in Czech language.
Dataset Details
Dataset Description
Curated by: Artificial Intelligence Center, FEE, CTU in Prague
Language(s) (NLP): Czech (cs, ces)
License: cc-by-nc-4.0
Dataset Sources
The data points in the dataset were collected from following sources:
MURI-IT - supernatural instructions, WikiHow, Reverse instructions… See the full description on the dataset page: https://huggingface.co/datasets/ctu-aic/cs_instruction_tuning_collection.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.en_instruction_tuning_collection
Dataset Card for English Instruction Tuning Collection
This dataset is a collection for instruction tuning of LLMs in English language. It is specifically designed as a counterpart of the Czech Instruction Tuning Collection.
Dataset Details
Dataset Description
Curated by: Artificial Intelligence Center, FEE, CTU in Prague
Language(s) (NLP): English (en)
License: cc-by-nc-4.0
Dataset Sources
The data points in the dataset were collected from… See the full description on the dataset page: https://huggingface.co/datasets/ctu-aic/en_instruction_tuning_collection.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Madhu348/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset-EDIT
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-EDIT.Turkish-Municipality-Instruction-Tuning-DatasetTrendyol-Cybersecurity-Instruction-Tuning-Dataset-EDIT
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/JEAPI/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-EDIT.mirror-Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Trendyol-Cybersecurity-Instruction-Tuning-Dataset.instruction_tuning_datasets
🧠 Persian Cultural Alignment Dataset for LLMs
This repository contains a high-quality, instruction-following dataset for cultural alignment of large language models (LLMs) in the Persian language. The dataset is curated using hybrid strategies that incorporate culturally grounded generation, multi-turn dialogues, translation, and augmentation methods, making it suitable for SFT, DPO, RLHF, and alignment evaluation.
📚 Dataset Overview
Domain
Methods Used… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/instruction_tuning_datasets.dynamics-of-instruction-tuning
DoIT: Dynamics of Instruction Tuning
DoIT is a collection of over 40k human-curated instruction-output pairs in Chinese. I created from https://huggingface.co/datasets/ChiyuSONG/dynamics-of-instruction-tuning.
It collects all data in dynamics-of-instruction-tuning/curated/full/*.json.
arsyra-instruction-tuning
🤖 ArSyra Instruction Tuning — Arabic LLM Fine-Tuning Data
Purpose-built data for fine-tuning Arabic LLMs with dialectal instruction pairs.
Dataset Summary
Arabic instruction-tuning data combining instruction-following pairs,
instruction descriptions, freeform responses, and quality control data.
3,300+ records designed for fine-tuning Arabic language models.
The hottest topic in NLP — instruction tuning — but for Arabic dialects.
Use this data for SFT (supervised… See the full description on the dataset page: https://huggingface.co/datasets/ArSyra/arsyra-instruction-tuning.
