datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sna-waxal-annotated-unlabeled
Shona WAXAL annotated-unlabeled checkpoint
This is a self-contained operational checkpoint for pseudo-labeling Shona ASR
data. It contains 90,253 conservatively segmented FLAC clips
(441.585 hours), but intentionally contains no transcripts.
Fields
transcription is intentionally empty.
speaker_id is an approximate source-blind EOM cluster or unknown;
speaker_clip_count is zero for unknown assignments.
gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.VietMed_unlabeled
unofficial mirror of VietMed (Vietnamese speech data in medical domain) unlabeled set
official announcement: https://arxiv.org/abs/2404.05659
official download: https://huggingface.co/datasets/leduckhai/VietMed
this repo contains the unlabeled set: 966h - 230k samples
i also gather the metadata: see info.csv
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-unlabeled.py
need to do: check misspelling, restore foreign words phonetised to… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_unlabeled.
