datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
personaplex-distill-conversations
PersonaPlex Distillation Conversation Dataset
Teacher-generated multi-turn conversation data for distilling/pruning NVIDIA PersonaPlex 7B
(a Moshi-architecture full-duplex speech-to-speech model).
What this is
Each sample is a real conversation rendered by the PersonaPlex teacher itself:
The student's turns are scripted (generated by Qwen3-8B) and voiced with Piper TTS
The teacher's (PersonaPlex's) responses are improvised live by the model — its real… See the full description on the dataset page: https://huggingface.co/datasets/niloy629/personaplex-distill-conversations.PersonaMix
PersonaMix
Controlled bilingual (Kazakh–English) benchmark for target-speaker ASR and target-presence detection on overlapping speech, released with Persona-ASR.
Four speakers (two female, two male) each read 11 scripted sentences in both Kazakh and English. Mixtures span 1–3 interfering speakers and SNRs of −3, 0, +3, +6 dB, under same-language (A) and cross-language (B) enrollment; the cross-language condition enrolls a speaker in one language and transcribes them in the… See the full description on the dataset page: https://huggingface.co/datasets/issai/PersonaMix.AItuber-Persona-Voices-JA
AItuber Persona Voices JA
195体のAITuberペルソナに対し、キャラクター設定に基づいた声質設計・セリフ生成・音声合成を行ったデータセットです。
概要
項目
値
ペルソナ数
195
総音声ファイル数
20,800 (参照音声195 + 発話20,600)
音声フォーマット
WAV, PCM 16-bit, 44.1kHz, mono
セリフカテゴリ
original, descriptive, emotional
言語
日本語
データ構造
各行は1つの音声ファイルに対応し、以下のカラムを持ちます:
カラム
型
説明
persona_id
string
ペルソナ識別子 (persona_000 ~ persona_194)
persona_index
int
ペルソナインデックス(元データセットの行位置に対応)
persona_name
string
キャラクター名
voice_description_ja… See the full description on the dataset page: https://huggingface.co/datasets/kizuna-intelligence/AItuber-Persona-Voices-JA.PersonalHub
Personal Hub: Exploring High-Expressiveness Speech Data through Spatio-Temporal Feature Integration and Model Fine-Tuning
Introduction
In this work, we present Personal Hub, a novel framework for mining and utilizing high-expressivity speech data by integrating spatio-temporal context with combinatorial attribute control. At the core of our approach lies a Speech Attribute Matrix, which enables annotators to systematically combine speaker-related features such as age… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/PersonalHub.uwb_atcosim_split_by_person_6_2_2
UWB-ATCC + ATCOSIM — combined ATC ASR dataset (ATCOSIM 6-2-2, gender-balanced)
A combined air-traffic-control ASR dataset built from Jzuluaga/uwb_atcc and Jzuluaga/atcosim_corpus, re-split with leakage-free (group-disjoint) train/validation/test splits. UWB-ATCC is split recording-disjoint (80/10/10) and ATCOSIM is split speaker-disjoint and gender-balanced, putting 1 male + 1 female speaker in each of validation and test (6 train / 2 / 2 speakers), with a fixed seed (42) and 16… See the full description on the dataset page: https://huggingface.co/datasets/mmonzel/uwb_atcosim_split_by_person_6_2_2.uwb_atcosim_split_by_person_8-1-1
UWB-ATCC + ATCOSIM — combined ATC ASR dataset (ATCOSIM 8-1-1)
A combined air-traffic-control ASR dataset built from Jzuluaga/uwb_atcc and Jzuluaga/atcosim_corpus, re-split with leakage-free (group-disjoint) train/validation/test splits. UWB-ATCC is split recording-disjoint (80/10/10) and ATCOSIM is split speaker-disjoint into 8 train / 1 validation / 1 test speakers, with a fixed seed (42) and 16 kHz audio. Columns: id, audio, text, segment_start_time, segment_end_time… See the full description on the dataset page: https://huggingface.co/datasets/mmonzel/uwb_atcosim_split_by_person_8-1-1.agentCourse_storagestorage for gaia question files
personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.Persona-Dialogue
Persona-Dialogue Dataset
Multi-turn persona-driven dialogue dataset with synthesized speech audio.
Overview
Total conversations: 21561
Total turns: 165871
Total audio duration: 498.0 hours
Audio format: WAV, mono, 24kHz
Language: English
Scenarios: 20
Scenarios
Scenario
Groups
Family life
7497
School classroom
2115
Company meeting
1776
Restaurant
1201
Travel group
983
Friends gathering
963
Library/Bookstore
962
Stadium/Sports game
915… See the full description on the dataset page: https://huggingface.co/datasets/Yifanfan/Persona-Dialogue.personal_mainpageExpresso-personal-projectpersona-voicesfamous-persons-20s
Dataset Card for "famous-persons-20s"
More Information needed
famous-persons-20s-raw
Languages Covered
The dataset includes audio and translation in the following languages:
English
Usage
from datasets import load_dataset
dataset = load_dataset("rohitdiwane/famous-persons-20s-raw")
CommonVoice-hindi-personal-projectpersona-voicespersonalized_data_clean_speaker5my_personal_tts_datasetmy_personal_tts_dataset_hipersonalized_data_clean_finalpersonalized_speaker3personalized_data_clean_speaker3personalized_data_clean_final-speaker3persona-voicepersonalpersonalized_data_clean_finalaugmentedpersonalized_speaker5Personagenspersonalized_data_cleanpersonalized_data_clean_finalaugmented-speaker3
