minspeech
minspeech
Minspeech (Southern Min)
Full train/dev/test transfer of the Minspeech Southern
Min (Taiwanese) speech corpus, staged from an internal COS copy. unlabeled is still just a
10-clip illustrative preview -- see "Splits" below.
Dataset Structure
id: official Minspeech utterance ID (S<series><episode><utterance> for train/dev/
test; source video ID for unlabeled).
audio: audio clip (16kHz mono WAV), spoken in Southern Min / Taiwanese.
mandarin: Mandarin Chinese… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/minspeech.minspeech
MinSpeech: Cleaned Multi-dialect Min-nan Dataset (Private)
Important Legal Notice & Copyright Status
This repository is a Private Research Fork of the MinSpeech corpus. It is maintained strictly for individual research purposes, specifically for fine-tuning Automatic Speech Recognition (ASR) and Speech-to-Text Translation (S2TT) models.
1. Ownership & Licensing
Annotations & Metadata: The transcriptions and segment metadata are derived from the MinSpeech… See the full description on the dataset page: https://huggingface.co/datasets/scbz/minspeech.
