datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.medical_fine_tuning_12MTCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.Instruction-tuning_DatasetsEmbedding-model-fine-tuning-datasetAgricultural_pests_and_diseases_instruction_tuning_datasymbolic-instruction-tuning
Symbolic Instruction Tuning
This is the offical repo to host the datasets used in the paper From Zero to Hero: Examining the Power of Symbolic Tasks in Instruction Tuning. The training code can be found in here.
TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.Visual-Extraction-Tuning-382K
Visual Extraction Tuning 382K
This repository contains the generated visual extraction tuning dataset from the paper Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models.
Project page: https://web.stanford.edu/~markendo/projects/downscaling_intelligence
Code: https://github.com/markendo/downscaling_intelligence
Overview
We provide the 382K examples generated using our visual extraction tuning data generation pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/markendo/Visual-Extraction-Tuning-382K.repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Volmarrs_Norse_Paganism_Fine-Tuning_Dataset_v1
Dataset Card for Volmarr's Norse Paganism Fine-Tuning Dataset v1
A comprehensive JSONL dataset of approximately 1000 high-quality training pairs designed for fine-tuning large language models on authentic Norse Paganism (Ásatrú/Heathenry) topics. Each pair features user queries about key concepts—such as introduction to Norse Paganism, cosmology, deities, creation myths, Ragnarok, religious practices, runes, sacred sites, and more—paired with detailed, lore-accurate responses in the… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/Volmarrs_Norse_Paganism_Fine-Tuning_Dataset_v1.maithili-instruction-tuningNCD_Instruct-Tuning_Medical_QA_Indonesian
IndoHealth-NLP Vol. 2: NCD Instruct-Tuning Medical QA (Sample)
📁 VIEW & DOWNLOAD SAMPLE FILES HERE
⚠️ DATASET LIMITATION NOTE:
This repository contains a FREE SAMPLE (200 rows) for evaluation purposes. To download the full, production-ready dataset containing 3,497 meticulously curated rows, please visit our official Gumroad page: [https://3929431511879.gumroad.com/l/IndoHealth-NLPVol2NCDInstruct-TuningMedicalQAIndonesian]
Dataset Summary
Building localized… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/NCD_Instruct-Tuning_Medical_QA_Indonesian.saferdecoding-fine-tuning
Dataset Card for SaferDecoding Fine Tuning Dataset
This dataset aims to fine-tune models in an attempt to defend against jailbreak attacks. It is an extension of SafeDecoding
Dataset Details
Dataset Description
The dataset generation process was adapted from SafeDecoding.
This dataset includes 252 original human-generated adversarial seed prompts, covering 18 harmful categories.
This dataset includes responses generated by Llama2, Vicuna, Dolphin, Falcon… See the full description on the dataset page: https://huggingface.co/datasets/aspear/saferdecoding-fine-tuning.medical_instruction_tuningGRAM-fine-tuning-65kThis is the dataset for Fine-tuning GRAM.
Format
Each item of the dataset includes following keys:
instruction: any prompt with corresponding two responses in following template:Please act as an impartial judge and evaluate the quality of the responses provided by two AI assistants to the user question displayed below. You should choose the assistant that follows the user's instructions and answers the user's question better.
Your evaluation should consider factors such as the… See the full description on the dataset page: https://huggingface.co/datasets/NiuTrans/GRAM-fine-tuning-65k.FibonacciAi-CODE-Fine_Tuning-DatasetDocument Version: 1.0.2 | Last Updated: 07/17/2026
Instruction-Tuning-with-GPT-4-RedPajama-Chat
Instruction Tuning with GPT 4 RedPajama-Chat
This dataset has been converted from the Instruction-Tuning-with-GPT-4 dataset for the purpose of fine-tuning the RedPajama-INCITE-Chat-3B-v1 model.
About Instruction-Tuning-with-GPT-4
English Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
Usage and License Notices
The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Instruction-Tuning-with-GPT-4-RedPajama-Chat.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/Soban1234/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.planner_instruction_tuning_2kBootstrap 2k Planner finetuning dataset for ReWOO.
It is a mixture of "correct" HotpotQA and TriviaQA task planning trajectories in ReWOO Framework.
Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/andycoco1128/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.fine_tuning_datraset_4_openaiportuguese-gpt3.5-fine-tuningTrendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ahmadkaab/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.fine_tuning_datraset_4_openaiua_gec_instruction_tuning
UA-GEC instruction tuning
This dataset contains prompts and expected outputs for the grammatical error
correction task in the Ukrainian language. It is based on the
CC-BY-4.0-licensed UA-GEC dataset. The
license of the original data is CC-BY-4.0.
This dataset contains 1,700 examples of fixing errors in long documents, and
~28,000 sentence-level examples.
The instructions ask to correct errors in the text. Sometimes the model outputs
the corrected text as is. At other times, it… See the full description on the dataset page: https://huggingface.co/datasets/osyvokon/ua_gec_instruction_tuning.fine_tuning521k-ja
fine_tuning521k-ja
This data is a dataset for fine-tuning the local language model (LLM). It consists of the translation of "ign_clean_instruct_dataset_500k" and "GPTeacher." Please feel free to use it. This dataset contains data such as Q&A, contextualized questions, role plays. Please contact us if you encounter any issues.
Since I'm not entirely clear on OpenAI's terms of service, please be cautious when using it for commercial purposes. There may be exceptions for… See the full description on the dataset page: https://huggingface.co/datasets/shumpei2525/fine_tuning521k-ja.
