datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Taiwanese-Minnan-Sutiau
Taiwanese-Minnan-Sutiau Dataset
The dataset consists of a curated collection of words that resemble tokens in Taiwanese Minnan (Taiwanese Hokkien), aimed at enhancing the recognition and processing of the language for various applications. Sourced from the Ministry of Education in Taiwan, this dataset serves as a valuable linguistic resource for researchers and developers engaged in language processing and recognition tasks.
Dataset Features
Source: Ministry of Education, Taiwan… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Sutiau.Taiwan-Tongues-ASR-CE-dataset-hokkien
Taiwan-Tongues-ASR-CE-dataset-hokkien
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.Taiwanese-Minnan-Example-Sentences
Taiwanese Minnan Example Sentences
The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems.
Dataset Features
Source: Ministry of Education, Taiwan (Sutian Resource Center)
Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.Taiwan-Tongues-ASR-CE-dataset-zhtw
Taiwan-Tongues-ASR-CE-dataset-zhtw
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.Taiwan-Tongues-ASR-CE-dataset-hakka
Taiwan-Tongues-ASR-CE-dataset-hakka
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hakka.taiwanspeechtaiwanspeech_mfaTaiwan-mandarin
Dataset Card for "Taiwan-mandarin"
More Information needed
zh-taiwanTaiwanese_ASRTaiwan-Tongues-ASR-CE-dataset-en
Taiwan-Tongues-ASR-CE-dataset-en
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-en.Taiwan_Mandarin_Speech_Data_by_Mobile_Phone_Reading
Dataset Card for Nexdata/Taiwan_Mandarin_Speech_Data_by_Mobile_Phone_Reading
Dataset Summary
This dataset is just a sample of Taiwan Mandarin Speech dataset(paid dataset) by mobile phone reading.The data collects 204 Taiwan residents with 450 sentences for each speaker. The recorded is rich in content, including economy, entertainment, news, spoken language, numbers, letters, etc., covering general scenes and human-computer interaction scenes. Manual transcription of text… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/Taiwan_Mandarin_Speech_Data_by_Mobile_Phone_Reading.zh-taiwan
Dataset Card for zh-taiwan
zh-taiwan 是一個繁體中文之語音資料集,總計 2,740 筆音頻(train 2,698 / val 14 / test 28),音訊取樣率為 16 kHz。每筆資料包含音頻、音頻長度、繁體中文文本與對應之正規化(簡體)文本,適用於繁體中文之語音合成(TTS)或語音辨識(ASR)模型訓練與評測。
本資料集原始來源為 ivanzhu109/zh-taiwan,本 repository 僅作為鏡像與格式整理之版本,原始著作權歸原作者所有。
Dataset Details
Dataset Description
本資料集提供 train / val / test 三個子集,每筆資料包含下列欄位:
audio:16 kHz WAV 音頻;
text:繁體中文文本,其中英文詞彙以全大寫形式保留(如 FIREFOXONANDROID、GOOGLE);
normalized_text:對應之簡體中文正規化文本,保留相同英文大寫形式;… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/zh-taiwan.taiwan-youtube-full-cleancv-taiwanesetaiwanese-youtube-ocr-robusttaiwanese-hokkien-ocr-testtaiwanese-youtube-ocr-first-5taiwanIndigenousSelectionRNN_Taiwantaiwanese-youtube-ocr-testTaiwanese_Hokkien_Tone_Recognition
