datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_instructions_122k_alpaca_stylesec-contracts-financial-extraction-instructions
S&P 500 SEC Financial Extraction Instructions
Dataset Summary
7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies:
Split
Examples
Filing Type
Description
train
3,430
Exhibit 10 + DEF 14A
Positive examples with validated outputs
corrective
4,253
Exhibit 10 + DEF 14A
Corrective, rescued, and negative examples
Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA,
Coding
Typescript coding,
Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.food-visual-instructions
Adapting Multimodal Large Language Models to Domains via Post-Training (EMNLP 2025)
This repos contains the food visual instructions for post-training MLLMs in our paper: On Domain-Specific Post-Training for Multimodal Large Language Models.
The main project page is: Adapt-MLLM-to-Domains
Data Information
Using our visual instruction synthesizer, we generate visual instruction tasks based on the image-caption pairs from extended Recipe1M+ dataset. These synthetic… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/food-visual-instructions.natural-instructionsPreprocessed version of Super-Natural-Instructions from https://github.com/allenai/natural-instructions/tree/master/splits. The same inputs may appear with different outputs, thus to avoid duplicate inputs, you can deduplicate by the id or the inputs field.
This is modified from https://huggingface.co/datasets/Muennighoff/natural-instructions
with a few improvements:
Adds positive/negative examples, outputs, explanations for each task, to
support different task definitions.
Adds an "eval"… See the full description on the dataset page: https://huggingface.co/datasets/jayelm/natural-instructions.instructions-pair-miningsmall-natural-instructionscore17-instructionsmosaic-instructions
Mosaic format for instructions dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-instructions.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-instructions
load it,
from streaming import LocalDataset… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-instructions.robust04-instructionsuner_llm_instructions
Dataset Card for Universal NER v1 in the Aya format
This dataset is a format conversion from its original v1 format into the Aya instruction format and it's released here under the same CC-BY-SA 4.0 license and conditions.
It contains data in multiple languages and this version is intended for multi-lingual LLM construction/tuning.
The dataset contains different subsets and their dev/test/train splits, depending on language.
Citation
If you utilize this dataset version… See the full description on the dataset page: https://huggingface.co/datasets/universalner/uner_llm_instructions.news21-instructionsreact-code-instructions
React Code Instructions
Popular Queries
Number of instructions by Model
Unnested Messages
Instructions Added Per Day
Dataset of Claude Artifact esque React Apps generated by Llama 3.1 70B, Llama 3.1 405B, and Deepseek Chat V3.
Examples
Virtual Fitness Trainer Website
LinkedIn Clone
iPhone Calculator
Chipotle Waitlist
Apple Store
pmc_llama_instructionsThis repo provides part of the dataset used for PMC-LLaMA-13B's instruction tuning.
Data
Size
Link
ChatDoctor
100K
https://www.yunxiangli.top/ChatDoctor/
MedQA
10.2K
https://huggingface.co/datasets/GBaker/MedQA-USMLE-4-options
MedMCQA
183K
https://huggingface.co/datasets/medmcqa
PubmedQA
211K
https://huggingface.co/datasets/pubmed_qa
LiveQA
635
https://huggingface.co/datasets/truehealth/liveqa
MedicationQA
690
https://huggingface.co/datasets/truehealth/medicationqa
UMLS… See the full description on the dataset page: https://huggingface.co/datasets/axiong/pmc_llama_instructions.russian_instructions_2June 10:
Почищены криво переведенные примеры кода
Добавлено >50000 человеческих примеров QA и инструкций
Обновленная версия русского датасета инструкций и QA.
Улучшения:
1. Увеличен размер с 40 мегабайт до 130 (60к сэмплов - 200к)
2. Улучшено качество перевода.
Структура датасета:
{
"sample":[
"Как я могу улучшить свою связь между телом и разумом?",
"Начните с разработки регулярной практики осознанности. 2. Обязательно практикуйте баланс на нескольких уровнях: физическом… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_instructions_2.law-instructions-dataset
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/law-instructions-dataset.open-korean-instructions4가지 한국어 챗봇 학습용 데이터셋을 합쳐놓았습니다. 이중 ShareGPT 데이터는 멀티턴으로 되어있습니다.
데이터 생성 및 합치는 코드는 https://github.com/HeegyuKim/open-korean-instructions 여기를 참고하세요
이름
#
타입
KoAlpaca v1.0
52K
싱글턴
KoAlpaca v1.1
21K
싱글턴
ShareGPT DeepL 번역
620K(싱글턴), 84K(멀티턴)
멀티턴, 싱글턴
OIG-small-chip2-ko
210K
싱글턴
Korquad-Chat
9.6K
멀티턴, 지식기반
모든 데이터는 포멧이 통일되어 있습니다. <sys>, <usr>, <bot> 세가지 토큰과 줄넘김으로 화자를 구분합니다.
korquad-chat 데이터의 경우, 유저와 봇이 서로를 호칭할 때는 <|bot|>, <|user|>로 되어있습니다.
{"source": "koalpaca-v1.0", "text":… See the full description on the dataset page: https://huggingface.co/datasets/heegyu/open-korean-instructions.python-code-instructions-85k
Python Code Instructions - 85K
Instruction-tuning dataset of Python functions paired with short natural-language instructions derived from repository docstrings.
What changed in this release
This release keeps the original public rows and format, but makes the dataset easier to use responsibly:
exact duplicate rows were removed again using normalized instruction + output hashing
deterministic train, validation, and test splits were added
the dataset card now documents… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/python-code-instructions-85k.russian_instructionsНовая версия: https://huggingface.co/datasets/Den4ikAI/russian_instructions_2
Русский датасет инструкций и QA.
Структура датасета:
{
"dialogue":[
"Как я могу улучшить свою связь между телом и разумом?",
"Начните с разработки регулярной практики осознанности. 2. Обязательно практикуйте баланс на нескольких уровнях: физическом, эмоциональном, умственном и духовном. 3. Свяжитесь с природой, когда это возможно - идите на прогулки или бегайте на улице, или просто сидите в парке и… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_instructions.rejected-agriculture-instructions-dataset
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-agriculture-instructions-dataset.agriculture-instructions-dataset
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/agriculture-instructions-dataset.unjudged-agriculture-instructions-dataset
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-agriculture-instructions-dataset.natural-instructions-samplepaper_instructions_300K-v1Loading will work as follows:
Existing behavior
# Loads the SFT dataset containing instruction, prompt, output
load_dataset("paperbd/paper_instructions_300K-v1")
Reasoning variant
# Loads reasoning subset containing instruction, prompt, reasoning, output
load_dataset(
"paperbd/paper_instructions_300K-v1",
"reasoning",
split="train",
)
Dataset Summary
This dataset contains synthetic supervised fine-tuning data generated from academic… See the full description on the dataset page: https://huggingface.co/datasets/paperbd/paper_instructions_300K-v1.lego-instructions-public
LEGO instruction PDF index
This public dataset contains one curated US-Letter or language-neutral visual
instruction PDF per current LEGO booklet. Explicit translated extras, obsolete
asset revisions, corrupt files, and duplicate file contents are excluded.
instruction-manifest.jsonl is the authoritative index. Each row records the
set number, year, set name, booklet position, retained LEGO asset identifier,
source URL, SHA-256 digest, byte size, and dataset-relative PDF path.… See the full description on the dataset page: https://huggingface.co/datasets/DonV1to/lego-instructions-public.unjudged-law-instructions-dataset
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-law-instructions-dataset.OrthoTryOn-Instructions
OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation
Model Introduction
We introduce OrthoTryOn, a unified and parameter-efficient framework for fashion image generation, designed to mitigate inter-task
interference in shared adaptation and enable high-quality virtual try-on, garment reconstruction, and pose transfer within a single model.
Its plug-and-play design can further extend to broader multi-task scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/Jerome-Young/OrthoTryOn-Instructions.unity-dev-instructions
Unity Developer Instructions
A comprehensive instruction-tuning dataset for Unity game development,
covering C# scripting, XR/VR development, physics, animation, rendering,
UI Toolkit, and performance optimization.
Dataset Summary
Split
Count
Train
46,483
Test
2,446
Total
48,929
Data Sources
| unity_docs | 40,496 |
| stackoverflow | 6,071 |
| github | 2,362 |
Source breakdown:
Source
Count
unity_docs
40,496
stackoverflow
6,071… See the full description on the dataset page: https://huggingface.co/datasets/vishnuOI/unity-dev-instructions.k8s-instructionsThis is a fork from https://huggingface.co/datasets/substratusai/k8s-instructions
LuauDev-instructions-SFT-preview
LuauDev-SFT-PREVIEW
THIS IS A PREVIEW VARIANT OF LUAUDEV.
non preview: Pinkstack/LuauDev-instructions-SFT-full
This is an SFT dataset meant for training Luau(Roblox's coding language) oriented large language models.
Once the full version would be out it would be the biggest Luau instruction-style dataset ever released.
These are the models which were used for data generation:
(no specific order)
DiffusionGemma 26B A4B
Deepseek v4 Flash 0731
Nemotron 3 Ultra 550B A55B
dots3 note… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstack/LuauDev-instructions-SFT-preview.
