datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
s2s-fr-finetuning
s2s-fr-finetuning
Corpus FR pour le finetuning speech-to-speech (Liquid-Audio / LFM2-Audio), construit par une
pipeline de prétraitement : VAD, ASR + alignement mot, segmentation aux frontières de mots,
filtrage qualité perceptuelle, normalisation de texte, déduplication.
Utilisation
from datasets import load_dataset
ds = load_dataset("baptistefrancois1/s2s-fr-finetuning", "common_voice_fr")
Un config HF par source d'origine : common_voice_fr, emilia_yodas_fr… See the full description on the dataset page: https://huggingface.co/datasets/baptistefrancois1/s2s-fr-finetuning.Urdu-Finetuning-Data-VibeVoice-Largeemotional-roleplay-finetuning-dataset
Artificial Voice Roleplay Dataset
67,491 fully-synthetic speech clips (~184 hours) pairing expressive role-play / character
voice-direction captions with generated audio, across German, English, Spanish, and French
(German-dominant). Rich in exaggerated fantasy/creature voices (orc, goblin, troll, ogre,
zombie, dragon, demon, witch, banshee, imp, fairy, gnome, robot, murloc, harpy, skeleton, ghost,
vampire …) and high-arousal emotional delivery (rage, fear, grief, menace).
Every… See the full description on the dataset page: https://huggingface.co/datasets/laion/emotional-roleplay-finetuning-dataset.Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning
Curated Arabic Speech Dataset for Seasme (from MCV17)
Dataset Description
This dataset is a curated and preprocessed version of the Arabic (ar) subset from Mozilla Common Voice (MCV) 17.0. It has been specifically prepared for fine-tuning conversational speech models, with a primary focus on the Seasme-CSM model architecture. The dataset consists of audio clips in WAV format (24kHz, mono) and their corresponding transcripts, along with integer speaker IDs.
The original… See the full description on the dataset page: https://huggingface.co/datasets/MAdel121/Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning.personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.OpenWhistle-Classification-Finetuning
OpenWhistle Classification Finetuning Dataset
OpenWhistleNeurIPS26/OpenWhistle-Classification-Finetuning is the public
classification finetuning dataset used for dolphin whistle identity
classification. It contains short whistle clips, whistle-level metadata,
fundamental-frequency tracks, rendered F0 spectrograms, and integer class
labels.
The main reviewer-facing subset is the balanced balanced config. It contains
six classes:
NSW_1 (label=0)
SW_Luna (label=1)
SW_Nana (label=2)… See the full description on the dataset page: https://huggingface.co/datasets/OpenWhistleNeurIPS26/OpenWhistle-Classification-Finetuning.fineTuningDatawhisper-finetuningtts-finetuningMDbA_FineTuningTTS_finetuning_datasetarticulationGAN_finetuning_data
Dataset Card for "articulationGAN_finetuning_data"
More Information needed
whisper-finetuning-for-aseeMDbA-FineTuningV2nova_finetuningwhisper_finetuningWhisper_FineTuning_KoWhisper_FineTuning_Sufine-tuning-Jewish
Dataset Card for "fine-tuning-Jewish"
More Information needed
Zeta-Voice-ID-Vits-Finetuningtecholution_noun_finetuning_v1whisper_finetuningwhisper-finetuning-audio-fileswhisper-finetuning-teststt_lowrank_finetuningOpenWhistle-Detection-Finetuning
OpenWhistle Detection Finetuning
Expert-annotated whistle-type detection dataset for OpenWhistle. Each example is a fixed-length 0.5 s audio window labeled with the whistle types present in that window; background/no-whistle windows are represented by an all-zero target vector.
Overview
Task: multi-label whistle-type detection on fixed-length audio windows
Target vector: label, with one binary decision per whistle type in the order SW_Neo, SW_Luna, SW_Nikita, SW_Nana… See the full description on the dataset page: https://huggingface.co/datasets/OpenWhistleNeurIPS26/OpenWhistle-Detection-Finetuning.fineTuningFine-Tuning-03parler-ttts-dataset-for-finetuningdatasets_finetuning_PhoWhisper
