datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PersonaMix
PersonaMix
Controlled bilingual (Kazakh–English) benchmark for target-speaker ASR and target-presence detection on overlapping speech, released with Persona-ASR.
Four speakers (two female, two male) each read 11 scripted sentences in both Kazakh and English. Mixtures span 1–3 interfering speakers and SNRs of −3, 0, +3, +6 dB, under same-language (A) and cross-language (B) enrollment; the cross-language condition enrolls a speaker in one language and transcribes them in the… See the full description on the dataset page: https://huggingface.co/datasets/issai/PersonaMix.personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.
