datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nadi2026-adi20-micro-25pct-knnvc
NADI 2026 ADI20-micro — kNN-VC augmented (4 target voices)
Voice-converted copy of the 25% stratified subset (seed 42) of
UBC-NLP/NADI_2026_ADI20_micro, made with kNN-VC
following Abdullah et al. 2025.
Configs: voice_01–voice_04, 16,757 rows each, train split only.
Validation/test audio is deliberately left natural.
Target voices: 4 Arabic speakers from Common Voice (~60s each), gender-balanced,
the same set used across all dialects.
column
meaning
audio
converted… See the full description on the dataset page: https://huggingface.co/datasets/nadi-task2/nadi2026-adi20-micro-25pct-knnvc.orpheus_tts_knn_vc_pathological
Combined Orpheus TTS 3B + Same-Speaker KNN Voice Conversion Dataset
Dataset Overview
This dataset contains fine-tuned Orpheus TTS 3B synthetic speech enhanced with same-speaker KNN voice conversion a
Enhancement Method: Orpheus TTS 3B LoRA fine-tuned synthetic speech → Same-speaker KNN voice conversion using real audio references from the same speaker
Dataset Statistics
Total Samples: 759
Total Duration: 2583.38 seconds (43.06 minutes)
Speakers: 8 speakers… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/orpheus_tts_knn_vc_pathological.spark_tts_knn_vc_pathological
Dataset Overview
Total Samples: 785
Total Duration: 3006.46 seconds (50.11 minutes)
Speakers: 8 speakers
Corpora: TORGO, UA-Speech, LibriSpeech
Sample Rate: 16kHz (KNN-VC output rate, native Spark compatibility)
Audio Format: WAV
sesame_tts_knn_vc_pathological
Dataset Overview
Total Samples: 759
Total Duration: 2892.82 seconds (48.21 minutes)
Speakers: 8 speakers
Corpora: TORGO, UA-Speech, LibriSpeech
Sample Rate: 16kHz (KNN-VC output rate)
Audio Format: WAV
pathological_knn_vc
Combined Same-Speaker KNN Voice Conversion Dataset
Dataset Statistics
Total Samples: 799
Total Duration: 0.4 hours
Speakers: 8
Corpora: TORGO, UA-Speech, LibriSpeech
Audio Format: 16kHz WAV
Baseline TTS: resproj007/baseline_orpheus_3b
Speaker Breakdown
Speaker
Name
Corpus
Condition
Gender
Samples
Duration
FC02
TORGO Healthy Female
TORGO
Healthy
Female
100
146.0s
M04
UA-Speech Male
UA-Speech
Dysarthric
Male
200
245.7s
M02
TORGO Dysarthric Male… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/pathological_knn_vc.
