datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
speechio_test
SpeechIO ASR Test Sets (parquet)
Parquet repackaging of the SpeechColab SpeechIO Mandarin ASR benchmark,
re-exported from yuekai/speechio (Lhotse cuts) into standard
HuggingFace parquet with embedded 16 kHz audio.
27 test sets: SPEECHIO_ASR_ZH00000 ... SPEECHIO_ASR_ZH00026, each a config with a single test split.
~43k utterances, ~66 hours total, evaluation only.
Columns
column
type
note
segment_id
string
utterance id
speaker
string
speaker id… See the full description on the dataset page: https://huggingface.co/datasets/yuekai/speechio_test.ePark_yue_du_shu_xie_pian_reading_writing
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_yue_du_shu_xie_pian_reading_writing
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_yue_du_shu_xie_pian_reading_writing.
