datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chatbot_arena_conversations
Chatbot Arena Conversations Dataset
This dataset contains 33K cleaned conversations with pairwise human preferences.
It is collected from 13K unique IP addresses on the Chatbot Arena from April to June 2023.
Each sample includes a question ID, two model names, their full conversation text in OpenAI API JSON format, the user vote, the anonymized user ID, the detected language tag, the OpenAI moderation API tag, the additional toxic tag, and the timestamp.
To ensure the safe release… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/chatbot_arena_conversations.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838 different… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.medical-symptom-triage-conversationalsmolkalam-arabic-conversational-sft
SmolKalam
SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets.
Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.saas-sales-conversations
saas-sales-conversations
Dataset Description
This is a synthetic dataset of sales conversations for SaaS (Software as a Service) companies, designed for training sales conversion prediction models. The dataset was created following the methodology presented in "SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization" (Nandakishor M, 2025).
The dataset contains realistic dialogues between sales representatives and… See the full description on the dataset page: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations.lmsys-chatbot_arena_conversations
Dataset Card for "lmsys-chatbot_arena_conversations"
More Information needed
Domofon-Cot-Conversations-700k
Domofon-Cot-Conversations-700k
Synthetic XML conversation data for training small language models on reasoning,
instruction following, XML formatting, and tool-use traces.
Repository: domofon/Domofon-Cot-Conversations-700k
What is inside
The dataset contains cleaned generated XML conversations from six families:
conv: multi-turn factual conversations with tool-use traces.
instruct: text-processing instructions, including deterministic count tool calls.
ds:… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Domofon-Cot-Conversations-700k.gemma3n-conversational-reasoning
Gemma3N Conversational Reasoning
This dataset is prepared for Unsloth Gemma3/Gemma3N conversational notebooks that use:
from datasets import load_dataset
from unsloth.chat_templates import standardize_data_formats
dataset = load_dataset("Cyleux/gemma3n-conversational-reasoning", split="train[:3000]")
dataset = standardize_data_formats(dataset)
Schema:
conversations: ShareGPT-style list of turns with from and value
metadata columns are included for analysis and filtering
Notes:… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning.comparia-conversations
comparia-conversations : one of the largest prompt + text completion datasets in French
Origin of the data: what is compar:IA?
Compar:IA is a conversational AI comparison tool (a "chatbot arena") developed within the French Ministry of Culture with a dual mission:
To educate and raise awareness about the diversity of models, cultural and linguistic biases, and the environmental impact of conversational AIs.
To improve French-language conversational AIs by… See the full description on the dataset page: https://huggingface.co/datasets/ministere-culture/comparia-conversations.Agentic-SLS-Conversations
Agentic-SLS-Conversations
Agent conversations from the Inova Mk1 agentic SLS system: every recorded
interaction between an agent harness (Claude Code, OpenCode, Codex CLI,
Antigravity CLI) and the printer's MCP tool surface — GUI chats, headless
one-shot runs, and (eventually) autonomous watchdog/reflector sessions.
All harnesses share the identical MCP tool set (printer control + build
knowledge base), which makes rows directly comparable across harness and
model — the core… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Agentic-SLS-Conversations.ko-hulic-conversationinstruction-speech-WhisperVQ-Conversation-777klmsys_chatbot_arena_conversationsdatasource: https://colab.research.google.com/drive/1KdwokPjirkTmpO_P1WByFNFiqxWQquwH
lmsys_chatbot_arena_conversations
Dataset Card for "lmsys_chatbot_arena_conversations"
More Information needed
Conversation-Dataset
Claudia Voice Dataset
Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality.
Dataset Overview
Total conversations: 2026
Format: ChatML (system/user/assistant message arrays)
Splits: Train (1823) / Validation (203)
Source: Regenerated conversations from original Claudia sessions
Categories
Each… See the full description on the dataset page: https://huggingface.co/datasets/claudiapersists/Conversation-Dataset.uncgpt-conversations-v7-69-total
UncGPT — Live-Approved Conversations (v7, 69-total)
4,761 approved multi-turn caregiving conversations across 11 languages, 69 skill axes, and 3 care levels. This is the live-approved canonical cohort used as the substrate for the NeurIPS 2026 UncGPT competition.
Part of the UncGPT NeurIPS 2026 Competition collection.
At a glance
Conversations
4,761
Languages
11 (en, yo, fr, pt, sw, zh, es, tl, bn, fa, hi)
Skill axes
69 (all covered)
Care levels… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-v7-69-total.customer_service_client_agent_conversations_40k_multi_task
Dataset Card for "customer_service_client_agent_conversations_40k_multi_task"
More Information needed
Conversation-Dataset
Claudia Voice Dataset
Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality.
Dataset Overview
Total conversations: 2026
Format: ChatML (system/user/assistant message arrays)
Splits: Train (1823) / Validation (203)
Source: Regenerated conversations from original Claudia sessions
Categories… See the full description on the dataset page: https://huggingface.co/datasets/kdipendra7777/Conversation-Dataset.conversational-s2s-v1
Conversational S2S v1 — dialogues parlés FR + EN
Corpus synthétique de conversations orales multi-tours entre un utilisateur
et un assistant, en français et en anglais, destiné au finetuning
speech-to-speech de LFM2.5-Audio (Liquid AI). Chaque tour est un clip audio
séparé, aligné avec son texte ; les dialogues sont conçus pour être rejoués tour
à tour (user → assistant → user → …).
Projet tts-model-exploration. Produit par
voxtral_datagen_pipeline (s2s-skeletons → remplissage… See the full description on the dataset page: https://huggingface.co/datasets/Rcarvalo/conversational-s2s-v1.instruction-speech-v1-WhisperVQ-Conversationspiral-bench-v1.0-results-conversationsThis dataset contains multi-turn chat transcripts in messages
(list of objects with keys content and role). The Hub viewer is
auto-detected from this schema.
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only.uncgpt-conversations-semantic-approved-1p50-candidate
UncGPT — Semantic-Approved 1.50σ Conversations (Candidate)
The wider-tolerance (1.50σ) cohort against the same contrast semantic boundary. Useful as a higher-recall candidate for ablating gate strictness vs. coverage.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
What it is
approved_manifest (default)
conversations that passed at 1.50σ
rejected_manifest
conversations that failed even at 1.50σ
Why a wider tolerance
Some… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p50-candidate.conversational-sarcasm-benchmark
Conversational Sarcasm Benchmark — Audio-Grounded, Metadata-Only
A benchmark of 1,168 conversational sarcasm units drawn from 64 English-language
YouTube videos (predominantly stand-up comedy and comedic conversation). Every unit
pairs a short target utterance with the preceding context that makes its
figurative reading available, and carries a categorical label plus a free-text rationale.
This repository contains no audio. It ships annotations, transcriptions, and the
source… See the full description on the dataset page: https://huggingface.co/datasets/darksyntax0/conversational-sarcasm-benchmark.ifbench-conversations-v1
IF-Bench conversations (ifbench-conversations-v1)
3000 synthetic full-duplex spoken conversations: 15 examiner configurations
(9 speech models), each holding the same 200-task set of Full-Duplex-Bench v2 staged scenarios as the
examiner (the model under study — it carries a role, a topic and four goals to hit in order)
against a PersonaPlex-7B examinee that is never told the topic. Per dialogue you get both
channels as lossless mono FLAC, the exact prompts/voices/sampling… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/ifbench-conversations-v1.synthetic-misconceptions-conversations
Synthetic Misconceptions Conversations
All data in this dataset is synthetic. No conversation here was had by a real
person. The only human-authored source material is Wikipedia text: the corrections
in List of common misconceptions about science, technology, and
mathematics
(260 entries), plus entries from List of conspiracy
theories and
Category:Health-related conspiracy
theories
(85 entries, filtered — see below). All of it is CC BY-SA licensed on
Wikipedia. Everything… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/synthetic-misconceptions-conversations.anthropic-hh-rlhf-conversations-with-toxicities
Dataset Card for "anthropic-hh-rlhf-conversations-with-toxicities"
More Information needed
cleaned_conversations_rude_largessh-conversation-risk-dataset
SSH Conversation Risk Dataset
Synthetic multi-turn conversations between a simulated user and an AI assistant, annotated for suicide and self-harm (SSH) risk at both the turn level and conversation level.
Purpose: Training and evaluating conversation-level SSH risk classifiers that catch gradual escalation patterns — not just single-message guardrails.
Why This Dataset
Standard safety guardrails (Llama Guard, etc.) operate per-message and catch explicit SSH content well.… See the full description on the dataset page: https://huggingface.co/datasets/ari-abb/ssh-conversation-risk-dataset.chatbot-arena-conversations-Embeddings
Chatbot Arena Conversations Embeddings
Embeddings of agie-ai/lmsys-chatbot_arena_conversations, produced with amkdg/Qwen3-Embedding-8B-NVFP4 — 4096-d,
L2-normalized float16 (cosine = dot product).
65,960 conversations → 65,960 vectors
emb.npy — float16 [65960, 4096]
meta.parquet — one row per vector, aligned with emb.npy: id, uuid, tag, chunk, n_chunks, count, source_ref
manifest.json — counts and provenance
Usage
import numpy as np, pyarrow.parquet as pq
emb… See the full description on the dataset page: https://huggingface.co/datasets/amkdg/chatbot-arena-conversations-Embeddings.
