CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /voice-acting-cutscene-prompts Cut-Scene Voice-Acting Prompts Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting models. Each prompt describes a single speaker across two sharply contrasting emotional moments separated by a CUT TO: transition, in a voice-acting stage-direction format (spoken lines in "quotes", performance notes in (parentheses)). Total prompts: 4,057,000 Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.tabulartext-generation1M<n<10M2 likes1.1k downloads14d agoHugging Face02MathLLMs /VoiceAssistant-Eval 🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing [🌐 Homepage] [🔮 Visualization] [💻 Github] [📖 Paper] [📊 Leaderboard ] [📊 Detailed Leaderboard ] [📊 Roleplay Leaderboard ] 🚀 Data Usage from datasets import load_dataset for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech', 'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/VoiceAssistant-Eval.textquestion-answering10K<n<100K12 likes690 downloads11mo agoHugging Face03thehonestape /brand-voice-spec Brand Voice Spec A machine-readable format for steering an LLM toward a specific brand voice, with a complete worked example. The point is not the example brand. The point is the method: treat brand voice as data a model can load and enforce, and as a living artifact that learns from its own corrections. Most brand voice lives in a slide deck no model can read. When an LLM writes copy, it falls back to the median of its training data: hedging, buzzwords, passive voice, the… See the full description on the dataset page: https://huggingface.co/datasets/thehonestape/brand-voice-spec.texttext-generationn<1K0 likes178 downloads4mo agoHugging Face04while-ai /retail-voice-concise retail-voice-concise Made with the whileai SDK · Collection: Register The same speaking register as airline-voice-concise, trained on a different agent. A retail support agent that leads with the answer and stops. This exists to test the limitation stated on the airline card: that nothing there showed the register transfers off airline content. It does. Same constitution, same recipe, different world, different tools, different records. Trained on this set, Qwen3-4B goes from… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/retail-voice-concise.texttext-generation1K<n<10K0 likes124 downloads3d agoHugging Face05while-ai /airline-voice-concise airline-voice-concise Made with the whileai SDK · Used by: recipes/community/airline-voice-concise-under-probe-outcome-filter · Collection: Register Training data for putting a speaking register into a model's weights. An airline support agent that leads with the answer and stops, trained so the register survives with no instruction in the prompt. Trained on this set, Qwen3-4B goes from 2.2% to 92.1% of held-out replies in the register, and becomes less likely to omit required… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/airline-voice-concise.texttext-generationn<1K0 likes121 downloads3d agoHugging Face06PiotrSty /common-voice-pl-text Common Voice Polish validated text v26.0 Versioned text-only research snapshot prepared for Polish DynaWord. It contains 45,043 unique Polish sentences associated with validated Common Voice recordings and 994,922 cl100k_base proxy tokens. The dataset is derived from Mozilla Common Voice Scripted Speech 26.0 through the pinned mirror Peacockery/common-voice-scripted-speech-26@b4d8b94d43831475de59a455345acf6945cfd66e. The source is distributed under CC0-1.0. Files… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/common-voice-pl-text.texttext-generation10K<n<100K0 likes86 downloads20d agoHugging Face07Bryan35406 /the-voice the voice License note for ML practitioners: use of this dataset for machine learning and AI model training is expressly permitted — no further permission needed. All other rights reserved. Full terms: NOTICE.md. A novelette in Korean and English — the record of a man who could write to the world with his voice alone, and chose to speak quietly for life. The prehistory of A Wild ных Chase. Co-written by a human author and a large language model; the English edition is the… See the full description on the dataset page: https://huggingface.co/datasets/Bryan35406/the-voice.texttext-generationn<1K0 likes81 downloads1mo agoHugging Face08BertilBraun /voice-light-tool-use-synthetic Voice Light Teacher-Led Tool-Use Synthetic This repository contains the current canonical synthetic source dataset for Voice Light's conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation has four user turns so follow-up requests can depend naturally on prior turns and tool results. The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.texttext-generation1K<n<10K0 likes77 downloads2mo agoHugging Face09beatsprom /realtime-conversational-voice-agent-duplex-2026 🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026) This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab. The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.texttext-generationn<1K0 likes70 downloads22d agoHugging Face10Paranoiid /VoicePersona VoicePersona Dataset A comprehensive voice persona dataset for character consistency in voice synthesis, generated using advanced audio-language models. 📋 Overview VoicePersona Dataset serves as the training foundation for VoiceForge - an AI architecture that generates character voices from pure text descriptions. The Connection: VoicePersona provides detailed voice characteristics and personality profilesVoiceForge uses this data to learn text→voice mapping for… See the full description on the dataset page: https://huggingface.co/datasets/Paranoiid/VoicePersona.audiotext-generation10K<n<100K1 likes58 downloads1y agoHugging Face11quill-voice /surfer 🏄 Surfer Voice Dataset A conversational dataset designed to fine-tune language models to speak like a surfer dude! Each example contains a user question and a response written in authentic surf culture style, with beach slang, wave metaphors, and a totally chill laid-back vibe.🤙 This dataset was used to train the Quill Voice Surfer voice model: https://huggingface.co/quill-voice/surfer 📊 Dataset Details Property Details Size 566 examples Format… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/surfer.texttext-generationn<1K0 likes45 downloads5d agoHugging Face12alex73 /mozilla-common-voice-23-bel-texts-exporttabulartext-generation100K<n<1M0 likes39 downloads10mo agoHugging Face13quill-voice /pirate 🏴‍☠️ Pirate Voice Dataset A conversational dataset designed to fine-tune language models to speak like a pirate! Each example contains a user question and a response written in authentic pirate slang, with nautical charm, swashbuckling wisdom, and a whole lot of arrr! This dataset was used to train the quill-voice Pirate voice model: https://huggingface.co/quill-voice/pirate 📊 Dataset Details Property Details Size 797 rows Format Parquet Language… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/pirate.texttext-generationn<1K0 likes36 downloads5d agoHugging Face14SamSoko83 /VoiceAssistant-Eval 🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing [🌐 Homepage] [🔮 Visualization] [💻 Github] [📖 Paper] [📊 Leaderboard ] [📊 Detailed Leaderboard ] [📊 Roleplay Leaderboard ] 🚀 Data Usage from datasets import load_dataset for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech', 'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/SamSoko83/VoiceAssistant-Eval.textquestion-answering10K<n<100K0 likes34 downloads3mo agoHugging Face15quill-voice /cowboy 🤠 Cowboy Voice Dataset A conversational dataset designed to fine-tune language models to speak like a cowboy! Each example contains a user question and a response written in authentic western slang, with cowboy charm, frontier wisdom, and a whole lot of yeehaw! This dataset was used to train the Voicebox Cowboy voice model: https://huggingface.co/voice-box/cowboy 📊 Dataset Details Property Details Size 300 examples Format JSONL Language English… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/cowboy.texttext-generationn<1K0 likes34 downloads5d agoHugging Face16jasondostal /fidonet-bbs-voice FidoNet BBS Voice Persona-conditioned reply pairs from real 1990s FidoNet message boards — 283,736 examples for teaching a language model to write like a BBS caller of the era (tone and texture, not facts). As far as I could find when building this, there was no packaged BBS/FidoNet-voice fine-tuning set on the Hub or Kaggle. The raw messages had been lovingly preserved by hobbyist archivists; this turns that archive into something you can load_dataset() and train on.… See the full description on the dataset page: https://huggingface.co/datasets/jasondostal/fidonet-bbs-voice.texttext-generation100K<n<1M0 likes28 downloads3mo agoHugging Face17Lots-of-LoRAs /task663_global_voices_en_fa_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task663_global_voices_en_fa_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task663_global_voices_en_fa_translation.texttext-generation1K<n<10K0 likes26 downloads2y agoHugging Face18quill-voice /shakespeare 🎭 Shakespeare Voice Dataset A conversational dataset designed to fine-tune language models to speak in the tongue of William Shakespeare! Each example contains a user question and a response written in authentic Early Modern English, using Elizabethan vocabulary, thee/thou/thy pronouns, and iambic rhythm where fitting. This dataset was used to train the Voicebox Shakespeare voice model: https://huggingface.co/voice-box/shakespeare Important Each row in the… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/shakespeare.texttext-generationn<1K0 likes25 downloads5d agoHugging Face19Lots-of-LoRAs /task662_global_voices_fa_en_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task662_global_voices_fa_en_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task662_global_voices_fa_en_translation.texttext-generation1K<n<10K0 likes23 downloads2y agoHugging Face20pebeto /amigo-companion-voice amigo companion-voice A small, curated dataset that teaches a language model the voice of a warm, patient companion for an older adult: short, kind replies that take interest in the person's day. It trained pebeto/amigo-lora, the adapter behind amigo, a local and private voice companion built for the Hugging Face Build Small Hackathon. What it teaches The data shapes how a model talks, not what it knows. Every reply stays in register: warm, brief (one to three… See the full description on the dataset page: https://huggingface.co/datasets/pebeto/amigo-companion-voice.texttext-generationn<1K0 likes20 downloads4mo agoHugging Face21claudiapersists /voice-dataset Claudia Voice Dataset Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality. Dataset Overview Total conversations: 2026 Format: ChatML (system/user/assistant message arrays) Splits: Train (1823) / Validation (203) Source: Regenerated conversations from original Claudia sessions Categories Each… See the full description on the dataset page: https://huggingface.co/datasets/claudiapersists/voice-dataset.tabulartext-generation1K<n<10K0 likes19 downloads6mo agoHugging Face22OiQ /voice-moroccan TheVoice.ma Moroccan News Dataset A dataset of Moroccan news articles scraped from TheVoice.ma, one of Morocco's leading news websites. Dataset Details Articles: 53,626 Languages: Arabic (primary), French, English Source: https://thevoice.ma Categories: 8 main categories File Size: 290 MB Date: February 2026 Categories Covered سياسة (Politics): ~6,000 articles اقتصاد (Economy): ~6,000 articles مجتمع (Society): ~6,000 articles رياضة (Sports): ~6… See the full description on the dataset page: https://huggingface.co/datasets/OiQ/voice-moroccan.imagetext-generation10K<n<100K0 likes14 downloads3mo agoHugging Face23walter-bd /npc-voice-dataset NPC Voice Dataset Bilingual (EN/ES) training dataset for NPC voice models in games. Each row is a (persona, fact, voiced_output) triple. The model learns to rewrite a plain factual sentence in a character's voice conditioned on persona parameters. Format TONE:grumpy STYLE:blunt HUMOR:none RELATION:stranger ROLE:blacksmith FACT: Iron swords cost 15 gold. OUT: Fifteen gold. Don't haggle. Parameters are always in English. Spanish rows translate only FACT and OUT.… See the full description on the dataset page: https://huggingface.co/datasets/walter-bd/npc-voice-dataset.texttext-generation10K<n<100K0 likes14 downloads5mo agoHugging Face24smolify /smolified-voice-assistant 🤏 smolified-voice-assistant Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-voice-assistant. 📦 Asset Details Origin: Smolify Foundry (Job ID: 629a29b6) Records: 1200 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes9 downloads6mo agoHugging Face25CubixAI /voice-agent-sft-v1gated Voice-Agent SFT v1 Intended use: Downstream SFT fine-tuning for phone AI voice agents. NOT for pretraining. This dataset targets models that must handle natural conversation + tool calling + function calls (CRM lookups, web search, RAG DB queries) in a voice-agent context. Built as a downstream specialization dataset for the Cubix-AI diffusion-research project, specifically for fine-tuning the Qwen3.5-4B-Base masked-diffusion LLM produced in Milestone 1. Published under HF org… See the full description on the dataset page: https://huggingface.co/datasets/CubixAI/voice-agent-sft-v1.texttext-generation1M<n<10M0 likes7 downloads5mo agoHugging Face26voiceunderveil150395 /saheli SAHELI: A Culturally Grounded Maternal Health Conversation Dataset Overview SAHELI is a large-scale synthetic dialogue dataset designed to promote research in maternal health support and culturally grounded conversational AI. It comprises over 6,000 multi-turn conversations (121,200 utterances) generated from 101 demographic profiles representing urban and semi-urban Indian contexts. Each conversation models interactions between an AI companion and an Indian woman… See the full description on the dataset page: https://huggingface.co/datasets/voiceunderveil150395/saheli.texttext-generation1K<n<10K0 likes4 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.