datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smolkalam-arabic-conversational-sft
SmolKalam
SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets.
Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.Persian-conversational-datasetpersian-conversational-datasetgemma3n-conversational-reasoning
Gemma3N Conversational Reasoning
This dataset is prepared for Unsloth Gemma3/Gemma3N conversational notebooks that use:
from datasets import load_dataset
from unsloth.chat_templates import standardize_data_formats
dataset = load_dataset("Cyleux/gemma3n-conversational-reasoning", split="train[:3000]")
dataset = standardize_data_formats(dataset)
Schema:
conversations: ShareGPT-style list of turns with from and value
metadata columns are included for analysis and filtering
Notes:… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning.TheArabicPile_Conversational
The Arabic Pile
Introduction:
The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Conversational.Multi-Turn-Conversational-SFTCreated by: DataCreator AI
Multi-Domain Multi-Turn Chat Conversations Dataset
A synthetic conversational dataset designed for LLM supervised fine-tuning and chatbot training.
The dataset contains multi-turn dialogues across multiple everyday domains such as travel, banking, health, programming, and customer interactions. Conversations are structured in OpenAI chat fine-tuning format, making the dataset directly usable in modern fine-tuning pipelines.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/DataCreatorAI/Multi-Turn-Conversational-SFT.childes-engUK-conversational-pairs
CHILDES Eng-UK Conversational Pairs
Curated naturalistic parent-child conversational pairs extracted from the
English-UK collection of CHILDES (MacWhinney, 2000), with a held-out test
set of 5 complete child histories that no model in the accompanying paper
has seen during training.
Dataset Summary
278,458 conversation pairs total across train, validation, and test
Train: 250,757 pairs from 2,784 transcripts
Validation: 13,197 pairs (in-distribution, sampled from… See the full description on the dataset page: https://huggingface.co/datasets/nshah-fbcs/childes-engUK-conversational-pairs.actuarial-conversational-dataset
Conversational Actuarial Dataset v0.1.0
Revolutionary Approach: Human First, Expert Second
This dataset transforms technical AI into conversational AI while maintaining domain expertise.
Dataset Composition
Total Examples: 461
56.6% Conversational: Natural dialogue, emotions, context
43.4% Technical: Actuarial with personality
Conversational Categories
Basic Interactions (56 examples)
Greetings and introductions
Small talk
Humor and… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-conversational-dataset.realtime-conversational-voice-agent-duplex-2026
🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026)
This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.mental_health_conversational_dataset
CREDIT: Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/ZahrizhalAli/mental_health_conversational_dataset.Conversational-Reasoning-Topical-Chat
Topical-Chat ShareGPT
This dataset is a modified version of the Conversational-Reasoning/Topical-Chat dataset, formatted in a ShareGPT-like structure for chatbot training. It comprises conversations between two individuals discussing news or Wikipedia articles, adapted for AI and user interaction. Please refer to the original dataset for comprehensive study details.
Dataset Description
Modifications
Key changes from the original dataset include:
AI and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/Conversational-Reasoning-Topical-Chat.conversational-finetuning-llama-format
Open Paws Conversational Finetuning Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Training Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.authentic-pre1930-sft-conversational
Pre-1930 Public Domain SFT Dataset
A supervised fine-tuning (SFT) dataset derived from 27 public-domain educational texts published before 1930, sourced from the Internet Archive. The texts span a wide range of 19th and early 20th century disciplines — natural science, history, law, philosophy, grammar, and more — and were written in a question-and-answer catechism format, making them naturally suited for instruction tuning.
Dataset Summary
Metric
Count… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/authentic-pre1930-sft-conversational.Bandori_Conversational_Benchmark_Action_Sequences
Codified Decision Tree (CDT) Dataset
Paper | GitHub
Codified Decision Trees (CDT) is a framework that induces executable and interpretable behavioral profiles for role-playing (RP) agents from narrative data. This dataset contains scene-action pairs derived from diverse storylines used to construct and validate these behavioral representations.
Introduction
Role-playing agents often rely on unstructured profiles that lead to brittle behavior. CDT represents behavioral… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Bandori_Conversational_Benchmark_Action_Sequences.SlimOrca-Dedup-trl-conversational-chatmlThis dataset is contains json formatted in TRL's conversational format as well as a chatml formatted text field.
Roblox_Luau_CoT_conversational_sharegpt_lqv1New version soon
This is a dataset based on Roblox/luau_corpus, with a sharegpt style, modified Thought and Output tokens, with a proper conversational style.
This highly experimental dataset is designed to help SLMs and LLMs handle reasoning with Luau Roblox code generation, it has the same style of tokens as Openo1
Ideally, after tuning your LLm with the Roblox/luau_corpus dataset, fine-tune it with another dataset like this one to create LLms similar to superthoughts by us, openo1, deepseek… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstack/Roblox_Luau_CoT_conversational_sharegpt_lqv1.Conversational_AOU_tutor_datasetgemma3n-conversational-reasoning-with-tools
Gemma3N Conversational Reasoning With Embedded Tool Traces
Prepared for Unsloth Gemma3/Gemma3N conversational notebooks that expect ShareGPT conversations.
Multi-turn conversations are preserved.
Reasoning blocks (<think>...</think>) are preserved.
Tool call traces are preserved by embedding them in assistant text as tags:
<tool_call ...>...</tool_call>
<tool_response ...>...</tool_response>
Use:
from datasets import load_dataset
from unsloth.chat_templates import… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning-with-tools.cole-conversational-corpus
Cole Multimodal Conversational Corpus
Real human-AI dialogue collected from a production AI assistant deployed across 7 messaging platforms. This is not synthetic data - these are real conversations where users send text, photos, voice messages, video, and documents alongside natural conversation.
What makes this dataset unique
Multimodal: Text, images, audio, video, and documents in conversation context - not isolated media files
Cross-channel: Same users move… See the full description on the dataset page: https://huggingface.co/datasets/huggingzetro/cole-conversational-corpus.TACTBench-Samples
TACTBench Demonstration Samples
This repository contains five full-context demonstration examples from
TACTBench. It does not contain the TACT training set or the remaining hidden
TACTBench evaluation set. The samples use the same full-history representation
as the benchmark evaluation and illustrate direct correction, error
explanation, guided revision, clarification checking, affective feedback, and
retry elicitation.
Data
data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.Pidgin-to-English-conversational-translations
Pidgin-to-English Translation Dataset (Sample)
Sample dataset: Nigerian Pidgin to English translation pairs for machine translation research
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The Pidgin-to-English Translation Dataset (Sample) is a conversational-style parallel corpus containing 122 translation pairs from Nigerian Pidgin English to Standard English. Created by Bytte AI through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-to-English-conversational-translations.gemma3n-conversational-reasoning-toolloop
Gemma3N Conversational Reasoning Tool-Loop
Gemma3N conversational dataset that preserves tool traces while avoiding training targets on tool responses.
Encoding:
Assistant emits tool calls: <tool_call ...>...</tool_call>
Tool outputs are user-side turns: <tool_response ...>...</tool_response>
This works with train_on_responses_only because user-side tool responses are masked from loss.
Use:
from datasets import load_dataset
from unsloth.chat_templates import… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning-toolloop.Conversational_pt_brDataset no estilo ShareGPT para treinamento de chatbots, com foco em conversas multi-turns sobre hardware.
prop-trading-qa-conversational-ai
Prop Trading Q&A Dataset for Conversational AI
Description
This dataset contains 200+ curated question-answer pairs covering the domain of proprietary (prop) trading firms. It is designed to serve as training and retrieval data for building AI assistants, chatbots, and educational tools focused on prop trading knowledge.
Each entry consists of a natural-language question paired with a detailed, factual answer. The data spans ten thematic categories ranging from… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/prop-trading-qa-conversational-ai.Conversational-HinglishConversational_AOU_tutor_datasetnemotron-gym-agentic-conversational-tool-use-pivot
laion/nemotron-gym-agentic-conversational-tool-use-pivot
Harbor task-binary dataset (96,965 tasks) converted from nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
(part of the nvidia/Nemotron-Post-Training-v3 collection).
Each row is a valid Harbor
task binary: columns path (str) and task_binary (gzip tar). Converted with the
OpenThoughts-Agent data.nemotron_gym framework.
Grading: Single-step: tool-call match (function_call) / LLM judge (message).
mental_health_conversational_dataset
CREDIT: Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/momed33/mental_health_conversational_dataset.
