datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
instruction-datasetThis is the blind eval dataset of high-quality, diverse, human-written instructions with demonstrations. We will be using this for step 3 evaluations in our RLHF pipeline.
Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.cpt_instruction_datasets
Instruction datasets
Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3.
Dataset creation
Datasets were created using two different techniques:
Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.Instruction-tuning_Datasetspao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset
An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub.
The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language.
The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.origen_dataset_instruction
OriGen: Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection
Introduction
OriGen is a fine-tuned lora model designed for Verilog code generation. It is trained on top of DeepSeek Coder 7B using datasets generated from code-to-code augmentation and self-reflection. The datasets can be found in the origen_dataset_instruction.
OriGen_Fix is a fine-tuned lora model designed for fixing syntax errors in Verilog code. It is trained based on OriGen… See the full description on the dataset page: https://huggingface.co/datasets/henryen/origen_dataset_instruction.ko-instruction-dataset
고품질 한국어 데이터셋
한국어로 이루어진 고품질 한국어 데이터셋 입니다.
WizardLM-2-8x22B 모델을 사용하여 WizardLM: Empowering Large Language Models to Follow Complex Instructions에서 소개된 방법으로 생성되었습니다.
@article{koinstructiondatasetcard,
title={CarrotAI/ko-instruction-dataset Card},
author={CarrotAI (L, GEUN)},
year={2024},
url = {https://huggingface.co/datasets/CarrotAI/ko-instruction-dataset}
}
full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.KrynexAI-Dataset-Flash-Instruction
🧠 KrynexAI Dataset
English | Русский
📌 Overview
KrynexAI Dataset is a high-quality, synthetically expanded collection of 10,000+ instruction-response pairs designed for fine-tuning Large Language Models (LLMs).
The dataset covers a wide range of topics including:
💻 Programming (Python, algorithms, data structures)
🤖 AI & Machine Learning (neural networks, transformers, LLMs)
🔭 Science (physics, cosmology, biology, neuroscience)
🧠 Philosophy & Psychology… See the full description on the dataset page: https://huggingface.co/datasets/KrynexLabs/KrynexAI-Dataset-Flash-Instruction.devops-cloud-instruction-dataset
DevOps & Cloud Infrastructure Dataset
Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure).
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.instruction_dataset
ProLLaMA Instruction Dataset
This repository contains the instruction dataset for ProLLaMA.
Paper
ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing
Code
GitHub Repository
Introduction
Recent advances in Protein Language Models (PLMs) have transformed protein engineering, yet unlike their counterparts in Natural Language Processing (NLP), current PLMs exhibit a fundamental limitation: they excel in either Protein… See the full description on the dataset page: https://huggingface.co/datasets/GreatCaptainNemo/instruction_dataset.telugu_instruction_dataset
Telugu Instruction Dataset — Luuka AI
Built by 10x Technologies — a curated Telugu-language instruction-tuning dataset for training the Luuka voice assistant.
Overview
This dataset contains 7,496 instruction–response pairs and 399 multi-turn conversations in Telugu, covering a broad range of natural voice assistant interactions. Every response is written entirely in Telugu script — no English characters appear in any response field.
Subset
Config Name
Pairs… See the full description on the dataset page: https://huggingface.co/datasets/10xtechnologieS/telugu_instruction_dataset.forum-instruction-tuning-dataset
Looksmaxxing Forum Dataset
A curated instruction-tuning dataset derived from a large looksmaxxing and aesthetic self-improvement forum,
containing high-density community knowledge on skincare, nutrition, supplementation, and appearance optimization.
Dataset Summary
This dataset was produced by scraping, parsing, cleaning, and LLM-filtering over 1.28 million raw forum posts
down to 77,417 high-quality instruction-response pairs using a multi-stage pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/bediss/forum-instruction-tuning-dataset.backend-api-instruction-dataset
Backend & API Development Dataset
Instruction dataset focused on RESTful API design, WebSocket real-time communication, microservices patterns, and API best practices.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Backend Api topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/backend-api-instruction-dataset.ArogyaAI-Medical-Instruction-Dataset
ArogyaAI Multimodal Medical Dataset 🏥
Dataset Description
The ArogyaAI Medical Dataset is a comprehensive, multimodal instruction-tuning and retrieval-augmented generation (RAG) dataset tailored specifically for the Indian healthcare context.
It was built to fine-tune the ArogyaAI LLaMA 3 8B model and BioLORD/BioBERT embeddings, bridging the gap between modern Allopathic medicine and traditional AYUSH systems (Ayurveda and Homeopathy).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Aman0026/ArogyaAI-Medical-Instruction-Dataset.turkish-poetry-instruction-dataset
Turkish Poetry Instruction Dataset
Türkçe şiir üretimi için Alpaca instruction formatında hazırlanmış fine-tuning dataseti.
Unsloth + Qwen LoRA eğitimi (3. ödev) için tasarlanmıştır.
Format
Her satır:
{
"instruction": "Kullanıcının şiir isteği",
"input": "",
"output": "Üretilmesi beklenen şiir metni"
}
Dosyalar
Dosya
Açıklama
train.jsonl
Eğitim (%90)
validation.jsonl
Doğrulama (%5)
test.jsonl
Test (%5)
Toplam ~4.961 örnek;… See the full description on the dataset page: https://huggingface.co/datasets/besmabakirci01/turkish-poetry-instruction-dataset.Instruction_recall_dataset
CanaryBench-PII
Frequency-aware canary injection benchmark for auditing memorization
in finetuned language models, built on the AI4Privacy PII reconstruction
task.
Dataset Description
This dataset is part of CanaryBench, a benchmark for evaluating
memorization in finetuned language models across repetition tiers
and privacy regimes.
Frequency tiers: 1×, 10×, 50×
PII types: EMAIL, PHONE
Member canaries: 770
Reference canaries: 1000
Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.ai-ml-instruction-dataset
AI/ML Engineering Instruction Dataset
Comprehensive instruction dataset covering machine learning concepts, PyTorch implementations, NLP with transformers, model evaluation, and feature engineering.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Ai Ml topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/ai-ml-instruction-dataset.Nepali_Education_Budget_ShareGPT_Instruction_Dataset
Nepali Education Budget — ShareGPT Instruction Dataset
merged_edu_sharegpt_ne_serial.jsonl
A Nepali-language, single-turn, fact-based Question–Answer dataset built from official Nepal government education budget statistics (Central Bureau of Statistics). The dataset is formatted in ShareGPT conversation style and is intended for instruction-tuning / fine-tuning language models to answer factual, numeric, statistics-grounded questions in pure Nepali (Devanagari script).… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepali_Education_Budget_ShareGPT_Instruction_Dataset.python-instruction-dataset
Python Developer Instruction Dataset
High-quality instruction-response pairs covering Python development best practices, async programming, decorators, type hints, and data manipulation with Pandas.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Python topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/python-instruction-dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/Soban1234/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ahmadkaab/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/andycoco1128/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.educational-math-instruction-dataset
Educational Math Instruction Dataset
Overview
This dataset contains instruction-style examples designed for fine-tuning large language models on educational mathematics tasks. The focus is on step-by-step reasoning, clear explanations, and pedagogically useful responses.
Dataset Structure
The dataset is provided in JSONL format. Each line contains a single instruction-response pair.
Intended Use
Fine-tuning LLMs for math tutoring
Educational… See the full description on the dataset page: https://huggingface.co/datasets/ushasree2001/educational-math-instruction-dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.turkish-finance-instruction-datasetThis dataset brings together three different sources to help research on financial question answering:
Synthetic data: Examples created automatically to cover more cases and add variety.
Manually curated data: Questions and answers written or labeled by people with domain knowledge.
Social media Q&A: Real question–answer pairs taken from finance discussions online.
Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Madhu348/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.pii_instruction_dataset_multilingual
