datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voice-acting-cutscene-prompts
Cut-Scene Voice-Acting Prompts
Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in a voice-acting stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
Total prompts: 4,057,000
Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.VoiceAssistant-Eval
🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
[🌐 Homepage]
[🔮 Visualization]
[💻 Github]
[📖 Paper]
[📊 Leaderboard ]
[📊 Detailed Leaderboard ]
[📊 Roleplay Leaderboard ]
🚀 Data Usage
from datasets import load_dataset
for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech',
'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/VoiceAssistant-Eval.brand-voice-spec
Brand Voice Spec
A machine-readable format for steering an LLM toward a specific brand voice, with a complete worked example. The point is not the example brand. The point is the method: treat brand voice as data a model can load and enforce, and as a living artifact that learns from its own corrections.
Most brand voice lives in a slide deck no model can read. When an LLM writes copy, it falls back to the median of its training data: hedging, buzzwords, passive voice, the… See the full description on the dataset page: https://huggingface.co/datasets/thehonestape/brand-voice-spec.retail-voice-concise
retail-voice-concise
Made with the whileai SDK · Collection: Register
The same speaking register as
airline-voice-concise,
trained on a different agent. A retail support agent that leads with the
answer and stops.
This exists to test the limitation stated on the airline card: that nothing
there showed the register transfers off airline content. It does. Same
constitution, same recipe, different world, different tools, different
records.
Trained on this set, Qwen3-4B goes from… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/retail-voice-concise.airline-voice-concise
airline-voice-concise
Made with the whileai SDK · Used by: recipes/community/airline-voice-concise-under-probe-outcome-filter · Collection: Register
Training data for putting a speaking register into a model's weights. An
airline support agent that leads with the answer and stops, trained so the
register survives with no instruction in the prompt.
Trained on this set, Qwen3-4B goes from 2.2% to 92.1% of held-out replies
in the register, and becomes less likely to omit required… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/airline-voice-concise.common-voice-pl-text
Common Voice Polish validated text v26.0
Versioned text-only research snapshot prepared for Polish DynaWord. It contains
45,043 unique Polish sentences associated with validated Common Voice
recordings and 994,922 cl100k_base proxy tokens.
The dataset is derived from Mozilla Common Voice Scripted Speech 26.0
through the pinned mirror Peacockery/common-voice-scripted-speech-26@b4d8b94d43831475de59a455345acf6945cfd66e. The source is
distributed under CC0-1.0.
Files… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/common-voice-pl-text.the-voice
the voice
License note for ML practitioners: use of this dataset for machine learning and AI model training is expressly permitted — no further permission needed. All other rights reserved. Full terms: NOTICE.md.
A novelette in Korean and English — the record of a man who could write to the world with his voice alone, and chose to speak quietly for life. The prehistory of A Wild ных Chase. Co-written by a human author and a large language model; the English edition is the… See the full description on the dataset page: https://huggingface.co/datasets/Bryan35406/the-voice.voice-light-tool-use-synthetic
Voice Light Teacher-Led Tool-Use Synthetic
This repository contains the current canonical synthetic source dataset for Voice Light's
conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English
conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation
has four user turns so follow-up requests can depend naturally on prior turns and tool results.
The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.realtime-conversational-voice-agent-duplex-2026
🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026)
This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.VoicePersona
VoicePersona Dataset
A comprehensive voice persona dataset for character consistency in voice synthesis, generated using advanced audio-language models.
📋 Overview
VoicePersona Dataset serves as the training foundation for VoiceForge - an AI architecture that generates character voices from pure text descriptions.
The Connection:
VoicePersona provides detailed voice characteristics and personality profilesVoiceForge uses this data to learn text→voice mapping for… See the full description on the dataset page: https://huggingface.co/datasets/Paranoiid/VoicePersona.surfer
🏄 Surfer Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a surfer dude!
Each example contains a user question and a response written in authentic surf culture style,
with beach slang, wave metaphors, and a totally chill laid-back vibe.🤙
This dataset was used to train the Quill Voice Surfer voice model:
https://huggingface.co/quill-voice/surfer
📊 Dataset Details
Property
Details
Size
566 examples
Format… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/surfer.mozilla-common-voice-23-bel-texts-exportpirate
🏴☠️ Pirate Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a pirate!
Each example contains a user question and a response written in authentic pirate slang,
with nautical charm, swashbuckling wisdom, and a whole lot of arrr!
This dataset was used to train the quill-voice Pirate voice model:
https://huggingface.co/quill-voice/pirate
📊 Dataset Details
Property
Details
Size
797 rows
Format
Parquet
Language… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/pirate.VoiceAssistant-Eval
🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
[🌐 Homepage]
[🔮 Visualization]
[💻 Github]
[📖 Paper]
[📊 Leaderboard ]
[📊 Detailed Leaderboard ]
[📊 Roleplay Leaderboard ]
🚀 Data Usage
from datasets import load_dataset
for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech',
'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/SamSoko83/VoiceAssistant-Eval.cowboy
🤠 Cowboy Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a cowboy!
Each example contains a user question and a response written in authentic western slang,
with cowboy charm, frontier wisdom, and a whole lot of yeehaw!
This dataset was used to train the Voicebox Cowboy voice model:
https://huggingface.co/voice-box/cowboy
📊 Dataset Details
Property
Details
Size
300 examples
Format
JSONL
Language
English… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/cowboy.fidonet-bbs-voice
FidoNet BBS Voice
Persona-conditioned reply pairs from real 1990s FidoNet message boards —
283,736 examples for teaching a language model to write like a BBS caller of the
era (tone and texture, not facts).
As far as I could find when building this, there was no packaged BBS/FidoNet-voice
fine-tuning set on the Hub or Kaggle. The raw messages had been lovingly preserved
by hobbyist archivists; this turns that archive into something you can
load_dataset() and train on.… See the full description on the dataset page: https://huggingface.co/datasets/jasondostal/fidonet-bbs-voice.task663_global_voices_en_fa_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task663_global_voices_en_fa_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task663_global_voices_en_fa_translation.shakespeare
🎭 Shakespeare Voice Dataset
A conversational dataset designed to fine-tune language models to speak in the
tongue of William Shakespeare! Each example contains a user question and a response
written in authentic Early Modern English, using Elizabethan vocabulary, thee/thou/thy
pronouns, and iambic rhythm where fitting.
This dataset was used to train the Voicebox Shakespeare voice model:
https://huggingface.co/voice-box/shakespeare
Important
Each row in the… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/shakespeare.task662_global_voices_fa_en_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task662_global_voices_fa_en_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task662_global_voices_fa_en_translation.amigo-companion-voice
amigo companion-voice
A small, curated dataset that teaches a language model the voice of a warm, patient companion for an older adult: short, kind replies that take interest in the person's day. It trained pebeto/amigo-lora, the adapter behind amigo, a local and private voice companion built for the Hugging Face Build Small Hackathon.
What it teaches
The data shapes how a model talks, not what it knows. Every reply stays in register: warm, brief (one to three… See the full description on the dataset page: https://huggingface.co/datasets/pebeto/amigo-companion-voice.voice-dataset
Claudia Voice Dataset
Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality.
Dataset Overview
Total conversations: 2026
Format: ChatML (system/user/assistant message arrays)
Splits: Train (1823) / Validation (203)
Source: Regenerated conversations from original Claudia sessions
Categories
Each… See the full description on the dataset page: https://huggingface.co/datasets/claudiapersists/voice-dataset.voice-moroccan
TheVoice.ma Moroccan News Dataset
A dataset of Moroccan news articles scraped from TheVoice.ma, one of Morocco's leading news websites.
Dataset Details
Articles: 53,626
Languages: Arabic (primary), French, English
Source: https://thevoice.ma
Categories: 8 main categories
File Size: 290 MB
Date: February 2026
Categories Covered
سياسة (Politics): ~6,000 articles
اقتصاد (Economy): ~6,000 articles
مجتمع (Society): ~6,000 articles
رياضة (Sports): ~6… See the full description on the dataset page: https://huggingface.co/datasets/OiQ/voice-moroccan.npc-voice-dataset
NPC Voice Dataset
Bilingual (EN/ES) training dataset for NPC voice models in games.
Each row is a (persona, fact, voiced_output) triple. The model learns to rewrite a plain factual sentence in a character's voice conditioned on persona parameters.
Format
TONE:grumpy STYLE:blunt HUMOR:none RELATION:stranger ROLE:blacksmith
FACT: Iron swords cost 15 gold.
OUT: Fifteen gold. Don't haggle.
Parameters are always in English. Spanish rows translate only FACT and OUT.… See the full description on the dataset page: https://huggingface.co/datasets/walter-bd/npc-voice-dataset.smolified-voice-assistant
🤏 smolified-voice-assistant
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-voice-assistant.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 629a29b6)
Records: 1200
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
voice-agent-sft-v1
Voice-Agent SFT v1
Intended use: Downstream SFT fine-tuning for phone AI voice agents. NOT for pretraining.
This dataset targets models that must handle natural conversation + tool calling + function
calls (CRM lookups, web search, RAG DB queries) in a voice-agent context.
Built as a downstream specialization dataset for the
Cubix-AI diffusion-research project,
specifically for fine-tuning the Qwen3.5-4B-Base masked-diffusion LLM produced in Milestone 1.
Published under HF org… See the full description on the dataset page: https://huggingface.co/datasets/CubixAI/voice-agent-sft-v1.saheli
SAHELI: A Culturally Grounded Maternal Health Conversation Dataset
Overview
SAHELI is a large-scale synthetic dialogue dataset designed to promote research in maternal health support and culturally grounded conversational AI. It comprises over 6,000 multi-turn conversations (121,200 utterances) generated from 101 demographic profiles representing urban and semi-urban Indian contexts.
Each conversation models interactions between an AI companion and an Indian woman… See the full description on the dataset page: https://huggingface.co/datasets/voiceunderveil150395/saheli.
