datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reviewed_sample
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.vietmed-reviewed
VietMed Reviewed
VietMed Reviewed is a reviewed Vietnamese medical speech dataset for automatic speech recognition.
This dataset is built from reviewed pseudo-labeled samples of the VietMed unlabeled subset. Each sample contains a WAV audio segment and a final reviewed transcript.
The dataset follows a simplified schema inspired by leduckhai/VietMed, with one split named reviewed.
Dataset Details
Task: Automatic Speech Recognition
Language: Vietnamese
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/phuvo05/vietmed-reviewed.
