datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
industrial-instruction-dataset
Industrial-Instruction Dataset
Industrial-Instruction provides benchmark and training-ready QA instances derived from industrial technical reports, designed to evaluate robustness under realistic retrieval conditions. Samples are grounded in retrieved evidence and include irrelevant retrieval, single-/multi-document support, and single-/multi-document answer settings.
Paper
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Parssky/industrial-instruction-dataset.cpt_instruction_datasets
Instruction datasets
Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3.
Dataset creation
Datasets were created using two different techniques:
Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.AddisGPT-Amharic-Instruction
AddisGPT-Amharic-Instruction
A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions.
796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.odia-instruction-dataset
Odia Instruction Following Dataset
Dataset Description
This is a comprehensive Odia language instruction-following dataset designed for training conversational AI models, chatbots, and instruction-following systems in Odia (ଓଡ଼ିଆ). The dataset contains high-quality instruction-response pairs that enable models to understand and follow instructions in the Odia language.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Total Records: 324,560
Format:… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-instruction-dataset.cybersecurity-controls-instructions
Cybersecurity Controls Instructions
Security control, incident response and risk management guidance from NIST Special Publications, turned into instruction-following examples.
Splits
split
rows
source documents
train
13,106
56
validation
4,840
18
test
5,697
18
Splits are held out by source document. Every chunk yields several
instruction rows, so a random row-level split would place the same passage in
train and test; whole documents are held… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/cybersecurity-controls-instructions.Instruction-280K-Turkish
Dataset Card for Instruction-280K-Turkish
Language: Turkish
Dataset Description
This repository contains a dataset for Turkish version of the Deepseek 1.5B model. The translation was performed using the Google translation model to ensure high-quality, accurate translation.
Dataset Details
Size: ≈280K
Translation tool: Google Translate
Data format: Prompt, Response
asha-instructions-hi-mr-v1
ASHA-Saathi Instructions (Hindi + Marathi) v1
A reusable Indic-medical instruction-tuning dataset for fine-tuning small language models to assist ASHA workers — India's ~1 million government-employed frontline community-health workers — in Hindi and Marathi.
Built as the training corpus for ombhojane/gemma-4-e2b-asha-it, submitted to the Gemma 4 Good Hackathon. Released independently so other researchers can fine-tune any small Indic-language LM on the same task.
Quick… See the full description on the dataset page: https://huggingface.co/datasets/ombhojane/asha-instructions-hi-mr-v1.
