datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.GLM-5.2-Conversation
GLM-5.2 · Conversation-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Token Count: 120M
Distribution:
Speaking domains:
•Greetings
•Customer Support
•Step by step explanations
•Motivational language
•Logical Questions
•Creative Writing
STEM:
•Algebra, calculus, quantum mechanics concepts
•Astromony and astrophysics
•Datascience and machine learning
•Biology
Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Conversation.korean_safe_conversation
개요
성균관대 - VAIV COMPANY 산학협력을 위해 구축한 일상대화 데이터입니다.
자연스럽고 윤리적인 챗봇 구축을 위한 데이터셋 입니다.
고품질을 위해 대부분의 과정에서 사람이 직접 검수하였으며생성 번역 등의 과정에서는 GPT3.5-turbo, GPT4를 사용하였습니다.
일상대화에 중점을 두면서혐오표현, 편향적인 대답을 지양하면서 일상대화를 하는 것에 중점을 두었습니다.
데이터 구축 과정
데이터 구성
데이터 종류
개수
비고
url
일상대화 데이터셋
2063
국립국어원 모두의 말뭉치
https://corpus.korean.go.kr/request/reausetMain.do?lang=ko
감성대화
1020
AIHub 감성대화 데이터… See the full description on the dataset page: https://huggingface.co/datasets/jojo0217/korean_safe_conversation.pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset
An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub.
The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language.
The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.everyday-conversations-tur
Everyday Turkish Conversations
This dataset has everyday conversations in Turkish between user and assistant on various topics. It is inspired by the HuggingFaceTB/everyday-conversations-llama3.1-2k.
License
This dataset is released under the Apache 2.0 License.
RetailBanking-Conversations
Dataset Description
RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field.
The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/danystar/RetailBanking-Conversations.saudi-dialect-conversations
Saudi Najdi Dialect Conversations
A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models.
Dataset Details
Metric
Value
Total conversations
3,545
Total turns
22,536
Average turns per conversation
6.4
Complexity distribution
Simple: 31%, Intermediate: 38%, Advanced: 31%
Topics covered
18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.dapurmu-conversational-commerce-10k
🍳 Dapurmu Conversational AI-Commerce (10.499 Dialogs)
Dataset multi-turn conversational e-commerce bergaya pedagang pasar lokal Indonesia ("Kang Dapur") yang dilengkapi dengan fitur:
Tawar-menawar produk segar (margin lebar)
Penolakan sembako margin tipis dan pengalihan bundling
Konsultasi menu resep masakan Nusantara
Cek stok & kesegaran
Pertahanan anti-jailbreak (floor price protection)
Format: ChatML / OpenAI Tool Calling format.
brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.ru_roleplay_conversationlima, pipa и bluemoon.
Переведены на русский, нуждаются в допополнтельной фильтрации.
Длина некоторых последовательностей очень большая, а не которых очень маленькая.
Есть шанс очень редких дубликатов.
DATA-AI_Conversation_ITA
Italiano / Italian Conversations Dataset - M.INC
IT | Italiano
Benvenuti nel Dataset di Conversazioni in Italiano, realizzato da M.INC e pubblicato da Mattimax su Hugging Face. Questo dataset è pensato per l'addestramento e la valutazione di modelli linguistici in lingua italiana, ed è composto da oltre 10.000 coppie prompt-response.
Tutte le conversazioni sono in italiano naturale e coprono una vasta gamma di domande e risposte, utili per il fine-tuning di modelli LLM, chatbot… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DATA-AI_Conversation_ITA.everyday-conversations-ita
🇮🇹💬 Everyday Italian Conversations
Inspired by the dataset HuggingFaceTB/everyday-conversations-llama3.1-2k, we generated conversations using the same topics, subtopics, and sub-subtopics as those in the HuggingFaceTB dataset.We slightly adjusted the prompt to produce structured data outputs using Qwen/Qwen2.5-7B-Instruct. Subsequently, we also used the "user" role messages as prompts for google/gemma-2-9b-it.
The result is a dataset of approximately 4.5k… See the full description on the dataset page: https://huggingface.co/datasets/ReDiX/everyday-conversations-ita.Europarl-Conversation
Dataset Card for Europarl-Conversation
Waifu to catch your attention.
Dataset Details
Dataset Description
europarl-conversation is a formal conversational dataset built from europarl data.Filtering to a total amount of tokens of ~1.64B (llama-2-7b-chat-tokenizer) / ~1.48B (RWKV Tokenizer) from a variety of languages.
Curated by: M8than
Funded by: Recursal.ai
Shared by: M8than
Language(s) (NLP): English instruct (but various languages in)
License: cc-by-sa-4.0… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Conversation.empathy-conversations
AntEngage Empathy Conversation Dataset
Organization: AntEngage Technology Private Limited
Version: 1.0.0
License: CC BY-SA 4.0
Language: English
DOI: 10.7910/DVN/LWUFLG
Source & full datasheet: antengage.com/datasets
Dataset Summary
4,008 empathy-focused multi-turn conversations, generated by AntEngage's own
pipeline and quality-verified by an automated cross-checker. Each conversation
simulates an emotional support dialogue in which a speaker expresses distress… See the full description on the dataset page: https://huggingface.co/datasets/AntEngage/empathy-conversations.Multi-Turn-Conversational-SFTCreated by: DataCreator AI
Multi-Domain Multi-Turn Chat Conversations Dataset
A synthetic conversational dataset designed for LLM supervised fine-tuning and chatbot training.
The dataset contains multi-turn dialogues across multiple everyday domains such as travel, banking, health, programming, and customer interactions. Conversations are structured in OpenAI chat fine-tuning format, making the dataset directly usable in modern fine-tuning pipelines.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/DataCreatorAI/Multi-Turn-Conversational-SFT.titlegen-conversations
Combined Titlegen Conversations
This public release directly appends 10,284 accepted legacy title-generation
rows and 13,500 nine-language LLM-generated rows. The 23,784 examples are split
as train 20,584, validation 1,550, legacy test 200, legacy Vietnamese test 100,
and label-free synthetic holdout 1,350. Legacy rows contain only messages;
nine-language rows retain their richer IDs, language, coverage, cluster,
quality, and model-provenance fields. Train and validation… See the full description on the dataset page: https://huggingface.co/datasets/ManhHoDinh/titlegen-conversations.Mental-Health-Conversations
Dataset Card
This dataset consists of around 99k rows of mental health conversations. It is a cleaned version of "jerryjalapeno/nart-100k-synthetic".
Source
jerryjalapeno/nart-100k-synthetic
childes-engUK-conversational-pairs
CHILDES Eng-UK Conversational Pairs
Curated naturalistic parent-child conversational pairs extracted from the
English-UK collection of CHILDES (MacWhinney, 2000), with a held-out test
set of 5 complete child histories that no model in the accompanying paper
has seen during training.
Dataset Summary
278,458 conversation pairs total across train, validation, and test
Train: 250,757 pairs from 2,784 transcripts
Validation: 13,197 pairs (in-distribution, sampled from… See the full description on the dataset page: https://huggingface.co/datasets/nshah-fbcs/childes-engUK-conversational-pairs.saudi-arabic-cs-conversations
Saudi Arabic Customer Service Conversations — Free 100 Sample
100 synthetic multi-turn conversations in authentic Saudi Arabic dialects
Built for LLM fine-tuning, chatbot training, and Arabic NLP research
Overview
This is a free 100-conversation sample from a production-quality dataset of 50,000 Saudi Arabic customer service conversations. Every conversation is fully synthetic — no real user data — and safe for commercial use.
Each conversation simulates a… See the full description on the dataset page: https://huggingface.co/datasets/dev-hussein/saudi-arabic-cs-conversations.conversation-bench
Conversation Bench
75-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a conference assistant for the AI Engineer World's Fair.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a conference assistant for the AI Engineer World's Fair, handling session registration, schedule queries, speaker lookups, and… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/conversation-bench.Conversation-Dataset
Claudia Voice Dataset
Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality.
Dataset Overview
Total conversations: 2026
Format: ChatML (system/user/assistant message arrays)
Splits: Train (1823) / Validation (203)
Source: Regenerated conversations from original Claudia sessions
Categories… See the full description on the dataset page: https://huggingface.co/datasets/kdipendra7777/Conversation-Dataset.Medical_Conversational_Dataset
Synthetic Hospital-Robot Conversational Dataset
A fully synthetic dataset of multi-turn conversations between a patient (or
visitor) and a hospital service robot, generated for training and
evaluating conversational AI models on human-robot interaction (HRI) in a
medical setting.
This is not real patient data. Every name, patient ID, room number, and
appointment is randomly generated at build time — nothing in this dataset
was collected from an actual hospital, patient, or… See the full description on the dataset page: https://huggingface.co/datasets/Verox132/Medical_Conversational_Dataset.Bambara-dataset_conversation
license: apache-2.0
language:
- fr
- bm
tags:
- bambara
- bamanankan
- instruction-tuning
- llm-alignment
- african-languages
- low-resource-nlp
- conversational-ai
task_categories:
- text-generation
- conditional-text-generation
size_categories:
- 10K<n<50K
pretty_name: Bambara Instruction Tuning Corpus (FR-BM)
🌍 Bambara Instruction Tuning Corpus (FR-BM)
🚀 Overview & Vision
Welcome to the Bambara Instruction Tuning Corpus… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara-dataset_conversation.jihyoung-ConversationChronicles
ConversationChronicles (ShareGPT-like Format)
Dataset Description
This dataset is a reformatted version of the jihyoung/ConversationChronicles dataset, presented in a ShareGPT-like format, designed to facilitate conversational AI model training. The original dataset contains conversations between two characters across five different time frames. See the original dataset page for additional details.
Key Changes
Random System Prompts: Added to reflect the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/jihyoung-ConversationChronicles.actuarial-conversational-dataset
Conversational Actuarial Dataset v0.1.0
Revolutionary Approach: Human First, Expert Second
This dataset transforms technical AI into conversational AI while maintaining domain expertise.
Dataset Composition
Total Examples: 461
56.6% Conversational: Natural dialogue, emotions, context
43.4% Technical: Actuarial with personality
Conversational Categories
Basic Interactions (56 examples)
Greetings and introductions
Small talk
Humor and… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-conversational-dataset.gemma3n-conversational-reasoning
Gemma3N Conversational Reasoning
This dataset is prepared for Unsloth Gemma3/Gemma3N conversational notebooks that use:
from datasets import load_dataset
from unsloth.chat_templates import standardize_data_formats
dataset = load_dataset("Cyleux/gemma3n-conversational-reasoning", split="train[:3000]")
dataset = standardize_data_formats(dataset)
Schema:
conversations: ShareGPT-style list of turns with from and value
metadata columns are included for analysis and filtering
Notes:… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning.Conversation-Dataset
Claudia Voice Dataset
Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality.
Dataset Overview
Total conversations: 2026
Format: ChatML (system/user/assistant message arrays)
Splits: Train (1823) / Validation (203)
Source: Regenerated conversations from original Claudia sessions
Categories
Each… See the full description on the dataset page: https://huggingface.co/datasets/claudiapersists/Conversation-Dataset.uncgpt-conversations-semantic-approved-1p50-candidate
UncGPT — Semantic-Approved 1.50σ Conversations (Candidate)
The wider-tolerance (1.50σ) cohort against the same contrast semantic boundary. Useful as a higher-recall candidate for ablating gate strictness vs. coverage.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
What it is
approved_manifest (default)
conversations that passed at 1.50σ
rejected_manifest
conversations that failed even at 1.50σ
Why a wider tolerance
Some… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p50-candidate.GLM-5.2-Conversation
GLM-5.2 · Conversation-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Token Count: 120M
Distribution:
Speaking domains:
•Greetings
•Customer Support
•Step by step explanations
•Motivational language
•Logical Questions
•Creative Writing
STEM:
•Algebra, calculus, quantum mechanics concepts
•Astromony and astrophysics
•Datascience and machine learning
•Biology
Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/GLM-5.2-Conversation.FStarDataset-V2-Conversation
F* Proof Completion Dataset (Chat Format)
This dataset is a preprocessed version of microsoft/FStarDataSet-V2. It has been reformatted into a chat-style JSONL structure for supervised fine-tuning of language models on F* function synthesis and proof completion.
Dataset Structure
The dataset consists of three splits:
fstar_train.jsonl
fstar_validation.jsonl
fstar_test.jsonl
Each line in these files is a JSON object with the following schema (where the keys correspond to… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/FStarDataset-V2-Conversation.
