datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voice-acting-cutscene-prompts
Cut-Scene Voice-Acting Prompts
Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in a voice-acting stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
Total prompts: 4,057,000
Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.VoiceAssistant-Eval
🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
[🌐 Homepage]
[🔮 Visualization]
[💻 Github]
[📖 Paper]
[📊 Leaderboard ]
[📊 Detailed Leaderboard ]
[📊 Roleplay Leaderboard ]
🚀 Data Usage
from datasets import load_dataset
for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech',
'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/VoiceAssistant-Eval.brand-voice-spec
Brand Voice Spec
A machine-readable format for steering an LLM toward a specific brand voice, with a complete worked example. The point is not the example brand. The point is the method: treat brand voice as data a model can load and enforce, and as a living artifact that learns from its own corrections.
Most brand voice lives in a slide deck no model can read. When an LLM writes copy, it falls back to the median of its training data: hedging, buzzwords, passive voice, the… See the full description on the dataset page: https://huggingface.co/datasets/thehonestape/brand-voice-spec.retail-voice-concise
retail-voice-concise
Made with the whileai SDK · Collection: Register
The same speaking register as
airline-voice-concise,
trained on a different agent. A retail support agent that leads with the
answer and stops.
This exists to test the limitation stated on the airline card: that nothing
there showed the register transfers off airline content. It does. Same
constitution, same recipe, different world, different tools, different
records.
Trained on this set, Qwen3-4B goes from… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/retail-voice-concise.airline-voice-concise
airline-voice-concise
Made with the whileai SDK · Used by: recipes/community/airline-voice-concise-under-probe-outcome-filter · Collection: Register
Training data for putting a speaking register into a model's weights. An
airline support agent that leads with the answer and stops, trained so the
register survives with no instruction in the prompt.
Trained on this set, Qwen3-4B goes from 2.2% to 92.1% of held-out replies
in the register, and becomes less likely to omit required… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/airline-voice-concise.common-voice-pl-text
Common Voice Polish validated text v26.0
Versioned text-only research snapshot prepared for Polish DynaWord. It contains
45,043 unique Polish sentences associated with validated Common Voice
recordings and 994,922 cl100k_base proxy tokens.
The dataset is derived from Mozilla Common Voice Scripted Speech 26.0
through the pinned mirror Peacockery/common-voice-scripted-speech-26@b4d8b94d43831475de59a455345acf6945cfd66e. The source is
distributed under CC0-1.0.
Files… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/common-voice-pl-text.the-voice
the voice
License note for ML practitioners: use of this dataset for machine learning and AI model training is expressly permitted — no further permission needed. All other rights reserved. Full terms: NOTICE.md.
A novelette in Korean and English — the record of a man who could write to the world with his voice alone, and chose to speak quietly for life. The prehistory of A Wild ных Chase. Co-written by a human author and a large language model; the English edition is the… See the full description on the dataset page: https://huggingface.co/datasets/Bryan35406/the-voice.voice-light-tool-use-synthetic
Voice Light Teacher-Led Tool-Use Synthetic
This repository contains the current canonical synthetic source dataset for Voice Light's
conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English
conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation
has four user turns so follow-up requests can depend naturally on prior turns and tool results.
The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.realtime-conversational-voice-agent-duplex-2026
🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026)
This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.VoicePersona
VoicePersona Dataset
A comprehensive voice persona dataset for character consistency in voice synthesis, generated using advanced audio-language models.
📋 Overview
VoicePersona Dataset serves as the training foundation for VoiceForge - an AI architecture that generates character voices from pure text descriptions.
The Connection:
VoicePersona provides detailed voice characteristics and personality profilesVoiceForge uses this data to learn text→voice mapping for… See the full description on the dataset page: https://huggingface.co/datasets/Paranoiid/VoicePersona.surfer
🏄 Surfer Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a surfer dude!
Each example contains a user question and a response written in authentic surf culture style,
with beach slang, wave metaphors, and a totally chill laid-back vibe.🤙
This dataset was used to train the Quill Voice Surfer voice model:
https://huggingface.co/quill-voice/surfer
📊 Dataset Details
Property
Details
Size
566 examples
Format… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/surfer.STT-Voice-Notes-Evals
STT Voice Note Evaluation
Author: Daniel RosehillDate Created: August 11, 2025Purpose: Comparative evaluation of Speech-to-Text (STT) services for voice note transcription
Overview
This dataset was created as part of ongoing work developing voice note transcription systems. It contains ground truth transcripts representing typical daily voice notes, recorded to evaluate and compare STT service accuracy across different content types.
Speaker Profile:
Single speaker… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/STT-Voice-Notes-Evals.mozilla-common-voice-23-bel-texts-exportVoiceAssistant-Eval
🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
[🌐 Homepage]
[🔮 Visualization]
[💻 Github]
[📖 Paper]
[📊 Leaderboard ]
[📊 Detailed Leaderboard ]
[📊 Roleplay Leaderboard ]
🚀 Data Usage
from datasets import load_dataset
for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech',
'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/SamSoko83/VoiceAssistant-Eval.pirate
🏴☠️ Pirate Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a pirate!
Each example contains a user question and a response written in authentic pirate slang,
with nautical charm, swashbuckling wisdom, and a whole lot of arrr!
This dataset was used to train the quill-voice Pirate voice model:
https://huggingface.co/quill-voice/pirate
📊 Dataset Details
Property
Details
Size
797 rows
Format
Parquet
Language… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/pirate.cowboy
🤠 Cowboy Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a cowboy!
Each example contains a user question and a response written in authentic western slang,
with cowboy charm, frontier wisdom, and a whole lot of yeehaw!
This dataset was used to train the Voicebox Cowboy voice model:
https://huggingface.co/voice-box/cowboy
📊 Dataset Details
Property
Details
Size
300 examples
Format
JSONL
Language
English… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/cowboy.task663_global_voices_en_fa_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task663_global_voices_en_fa_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task663_global_voices_en_fa_translation.fidonet-bbs-voice
FidoNet BBS Voice
Persona-conditioned reply pairs from real 1990s FidoNet message boards —
283,736 examples for teaching a language model to write like a BBS caller of the
era (tone and texture, not facts).
As far as I could find when building this, there was no packaged BBS/FidoNet-voice
fine-tuning set on the Hub or Kaggle. The raw messages had been lovingly preserved
by hobbyist archivists; this turns that archive into something you can
load_dataset() and train on.… See the full description on the dataset page: https://huggingface.co/datasets/jasondostal/fidonet-bbs-voice.shakespeare
🎭 Shakespeare Voice Dataset
A conversational dataset designed to fine-tune language models to speak in the
tongue of William Shakespeare! Each example contains a user question and a response
written in authentic Early Modern English, using Elizabethan vocabulary, thee/thou/thy
pronouns, and iambic rhythm where fitting.
This dataset was used to train the Voicebox Shakespeare voice model:
https://huggingface.co/voice-box/shakespeare
Important
Each row in the… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/shakespeare.common_voice_11_clean_tokenizedA cleaned and tokenized version of the English data from Mozilla Common Voice 11 dataset.
Cleaning steps:
Filtered on samples with >2 upvotes and <1 downvotes]
Removed non voice audio at start and end through pytorch VAD
Tokenization:
Audio tokenized through EnCodec by Meta
Using 24khz pre-trained model, and target bandwidth of 1.5
Represented in text as audio_token_0 - audio_token_1023
Prompts constructed as "text: <common voice transcript>\naudio: <audio tokens>"
Prompts tokenized with… See the full description on the dataset page: https://huggingface.co/datasets/anforsm/common_voice_11_clean_tokenized.task662_global_voices_fa_en_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task662_global_voices_fa_en_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task662_global_voices_fa_en_translation.hindi_voice_transcriptionsvoice-dataset
Claudia Voice Dataset
Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality.
Dataset Overview
Total conversations: 2026
Format: ChatML (system/user/assistant message arrays)
Splits: Train (1823) / Validation (203)
Source: Regenerated conversations from original Claudia sessions
Categories
Each… See the full description on the dataset page: https://huggingface.co/datasets/claudiapersists/voice-dataset.amigo-companion-voice
amigo companion-voice
A small, curated dataset that teaches a language model the voice of a warm, patient companion for an older adult: short, kind replies that take interest in the person's day. It trained pebeto/amigo-lora, the adapter behind amigo, a local and private voice companion built for the Hugging Face Build Small Hackathon.
What it teaches
The data shapes how a model talks, not what it knows. Every reply stays in register: warm, brief (one to three… See the full description on the dataset page: https://huggingface.co/datasets/pebeto/amigo-companion-voice.egyptian-voice-commands
Egyptian Voice Commands Dataset
This repository contains the Egyptian Arabic voice commands dataset used for training and evaluating the EgyptianAgent ASR and NLU models.
Dataset Structure
egyptian_voice_commands/
├── train.jsonl # Training data (665 examples)
├── eval.jsonl # Validation data (50 examples)
└── test.jsonl # Test data (102 examples)
egyptian_ui_navigation/
├── train.jsonl # Training data (50 examples)
└── test.jsonl # Test data… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/egyptian-voice-commands.npc-voice-dataset
NPC Voice Dataset
Bilingual (EN/ES) training dataset for NPC voice models in games.
Each row is a (persona, fact, voiced_output) triple. The model learns to rewrite a plain factual sentence in a character's voice conditioned on persona parameters.
Format
TONE:grumpy STYLE:blunt HUMOR:none RELATION:stranger ROLE:blacksmith
FACT: Iron swords cost 15 gold.
OUT: Fifteen gold. Don't haggle.
Parameters are always in English. Spanish rows translate only FACT and OUT.… See the full description on the dataset page: https://huggingface.co/datasets/walter-bd/npc-voice-dataset.voice-moroccan
TheVoice.ma Moroccan News Dataset
A dataset of Moroccan news articles scraped from TheVoice.ma, one of Morocco's leading news websites.
Dataset Details
Articles: 53,626
Languages: Arabic (primary), French, English
Source: https://thevoice.ma
Categories: 8 main categories
File Size: 290 MB
Date: February 2026
Categories Covered
سياسة (Politics): ~6,000 articles
اقتصاد (Economy): ~6,000 articles
مجتمع (Society): ~6,000 articles
رياضة (Sports): ~6… See the full description on the dataset page: https://huggingface.co/datasets/OiQ/voice-moroccan.smolified-voice-assistant
🤏 smolified-voice-assistant
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-voice-assistant.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 629a29b6)
Records: 1200
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
voice-agent-sft-v1
Voice-Agent SFT v1
Intended use: Downstream SFT fine-tuning for phone AI voice agents. NOT for pretraining.
This dataset targets models that must handle natural conversation + tool calling + function
calls (CRM lookups, web search, RAG DB queries) in a voice-agent context.
Built as a downstream specialization dataset for the
Cubix-AI diffusion-research project,
specifically for fine-tuning the Qwen3.5-4B-Base masked-diffusion LLM produced in Milestone 1.
Published under HF org… See the full description on the dataset page: https://huggingface.co/datasets/CubixAI/voice-agent-sft-v1.saheli
SAHELI: A Culturally Grounded Maternal Health Conversation Dataset
Overview
SAHELI is a large-scale synthetic dialogue dataset designed to promote research in maternal health support and culturally grounded conversational AI. It comprises over 6,000 multi-turn conversations (121,200 utterances) generated from 101 demographic profiles representing urban and semi-urban Indian contexts.
Each conversation models interactions between an AI companion and an Indian woman… See the full description on the dataset page: https://huggingface.co/datasets/voiceunderveil150395/saheli.
