datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sujet-Finance-Instruct-177k
Sujet Finance Dataset Overview
The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.turkish_instructionsNemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.SIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8B-Instructinvestopedia-instruction-tuning-dataset
Dataset Card for investopedia-instruction-tuning dataset
We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data
and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that
ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-instruction-tuning-dataset.QA_Instruction
𓅰 FINCH: CoT-Instruction Dataset for Korean Finance 𓅰
Overview
FINCH is a CoT-Instruction dataset rooting Korean-Financial tasks including: Multiple-Choice Question Answering (MCQA),
Extractive Question Answering (EQA), Binary Question Answering (BQA), Numerical Reasoning, Tabular Reasoning and Sentiment Analysis.
Additional details, research paper and further updates are coming! Stay Tuned.
Bangla-Instruct
Accepted in ACL Main 2025
TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
George Mason University, VA, USA
mraihan2@gmu.edu
If you find our work helpful, please consider citing our paper:
@inproceedings{raihan-zampieri-2025-tigerllm,
title = "{T}iger{LLM} - A Family of {B}angla Large Language Models",
author = "Raihan, Nishat and
Zampieri, Marcos",
editor = "Che, Wanxiang and
Nabende, Joyce and… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Instruct.Legal_Clause_InstructionsMyanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.Hinglish_Dataset_instruction_and_rawSIGNAL-Dataset-Hiddens-RefalMachine-RuadaptQwen2.5-7B-InstructChEMBL_Drug_Instruction_Tuning
Dataset Card for ChEMBL Drug Instruction Tuning
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/alxfgh/ChEMBL_Drug_Instruction_Tuning.marketing-instruct-4k
Dataset Card for marketing-instruct-4k
Dataset Details
Dataset Description
A curated instruction-tuning dataset of ~4,300 marketing copywriting
examples across five task types, built for the AutoScientist Challenge
2026 (Marketing category). Used to fine-tune Marketing-Mixtral-8x7B.
Key finding: this carefully curated dataset at its natural size
outperformed a 12,000-row version expanded via automated augmentation
(80% vs 58% win rate against the… See the full description on the dataset page: https://huggingface.co/datasets/suehuynh/marketing-instruct-4k.AddisGPT-Amharic-Instruction
AddisGPT-Amharic-Instruction
A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions.
796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.girt-instruct
GIRT-Instruct Corpus
Paper: https://arxiv.org/abs/2402.02632
A dataset in the format of pairs of instructions and corresponding outputs. GIRT-Instruct is constructed based on GIRT-Data, a dataset of IRTs.
We use both GIRT-Data metadata and the Zephyr-7B-Beta language model to generate the instructions
This dataset is used to train the GIRT-Model model.
Model: model
Space: space
Type
We have 4 different types in GIRT-Instruct. These types include:
default: This type… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/girt-instruct.PubChem_Drug_Instruction_Tuninginstructpoet-ar
Arabic Poetry IFT
Dataset Summary
Arabic Poetry IFT is a large-scale instruction-following dataset for Arabic poetry understanding and co-creation. It supports four task families: generation, continuation, revision/restoration, and multiple-choice analysis. The dataset covers Modern Standard Arabic (MSA) and four regional Arabic varieties used in the instruction layer: Gulf, Levantine, Nile Valley, and North African Arabic.
This release accompanies the ACL 2026 paper… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/instructpoet-ar.code_instructions45k_python_code_chinese_instruction
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
中文提示的代码数据集
其中提示部分通过调用GPT-4.0-turbo API翻译成中文
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/jean1/45k_python_code_chinese_instruction.Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only
Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only.safer-instruct
Safer-Instruct: Aligning Language Models with Automated Preference Data
This repository contains the dataset for the paper titled "Safer-Instruct: Aligning Language Models with Automated Preference Data". Check out our project website here!
Abstract
Reinforcement learning from human feedback (RLHF) is a vital strategy for enhancing model capability in language models. However, annotating preference data for RLHF is a resource-intensive and creativity-demanding process… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/safer-instruct.collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources:
Turkish Poems Cleaned
Turkish Reading Comprehension Question Answering Dataset
Stanford ALPaCA Cleaned Turkish Translated
Turkish Poems
Turkish Folk Song Lyrics
The data has been merged and processed for quality and consistency to create this dataset.
Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Adversarial-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only.Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only.instructional-dialogues-multilingual
Multilingual Instructional Dialogues (10-Language Dataset)
Multilingual Instructional Dialogues is a high-quality dataset of 100 structured, goal-oriented dialogues in 10 major world languages, created for training and fine-tuning AI assistants, chatbots, and instruction-tuned large language models.
Each dialogue simulates a clear, polite interaction where a user asks for guidance on how to perform a task, and the assistant responds with easy-to-follow steps. This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/Raftico/instructional-dialogues-multilingual.Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.tiny-instruct
tiny-instruct-v1
This dataset is collated from multiple other open-source datasets (de-duplicated). This has a total of ~6M rows each with an instruction and response (single-turn converstion).
Code Datasets:
CodeAlpaca_20K
CodeExercise-Python-27k
Evol-Instruct-Code-80k-v1
tiny-codes
Evol-instruction-66k
sciphi-python-textbook
programming_books_llama
WizardLM_evol_instruct_70k
Math Datasets:
MetaMathQA
arxiv-math-instruct-50k
MathInstruct
General… See the full description on the dataset page: https://huggingface.co/datasets/04RR/tiny-instruct.SIGNAL-Dataset-Hiddens-RefalMachine-RuadaptQwen2.5-14B-Instructturkish-reasoning-instructionsQwen2.5-7B-Instruct-em-eval
