datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moe-speech
Recomendation of OOPPEENN's 56697375616C4E6F76656C5F44617461736574
I recommend OOPPEENN/56697375616C4E6F76656C5F44617461736574 for Japanese voice corpus, which is:
Similar speech domain to this one (Japanese anime-style speech from Japanese Visual Novel), but
Huge amounts of audio compared to this dataset (600 hours for this, 10,000 hours for Galgame_Dataset!)
This dataset contains about 50 games, and Galgame_Dataset contains more than 500 games!
Contains true transcripts of each… See the full description on the dataset page: https://huggingface.co/datasets/litagin/moe-speech.College-Entrance-English-Examination-Listening-PartIf our dataset is useful for your research, please cite our work:
@article{li2025uni,
title={Uni-moe: Scaling unified multimodal llms with mixture of experts},
author={Li, Yunxin and Jiang, Shenyuan and Hu, Baotian and Wang, Longyue and Zhong, Wanqi and Luo, Wenhan and Ma, Lin and Zhang, Min},
journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
year={2025},
publisher={IEEE}
}
tat-moe
TAT-MOE
TAT (Taiwanese Across Taiwan) MOE 台文語音語料庫 -- speakers reading aloud example
sentences from the Ministry of Education's official Taiwanese dictionary (臺灣台語常用詞
辭典), recorded simultaneously on 6 different microphones per sentence.
⚠️ Known incompleteness -- this is NOT the full official TAT-MOE corpus. The official
release covers 328 train / 58 eval / 54 test speakers (86,072 / 16,357 / 15,962 sentences).
What is actually available in the source COS bucket -- and therefore… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tat-moe.moe-training
Xoron-Dev Multimodal MoE Dataset
This dataset is a high-scale, unified multimodal collection specifically engineered for training Mixture of Experts (MoE) models. It integrates text, audio, image, and video data into a single, cohesive training pipeline designed to foster cross-modal reasoning, creative generation, and agentic behavior.
🚀 Capabilities
By utilizing this dataset, models can be trained for:
Vision-Language: Image generation, high-fidelity editing… See the full description on the dataset page: https://huggingface.co/datasets/Backup-bdg/moe-training.tat-moe-subsetmoe-speech-plus
MoeSpeechPlus
MoeSpeechのデータセットに以下のデータ処理を行ったデータセット
追加情報
anime-whisperによる文字お越し
parakeetによる文字お越し
speechMOS (UTMOS v2 (T05)を使用)
音声の長さ
microsoft/deberta-v3-largeによる感情判定
litagin/anime_speech_emotion_classificationによる感情判定
Qwen/Qwen2-Audio-7B-Instructによる感情判定
{
"parakeet_jp_transcription": "昨夜からずっと気配を探られていたか。",
"anime_whisper_transcription": "昨夜からずっと気配を探られていたか…",
"duration": 3.722018140589569,
"speechMOS": 2.294760227203369,
"DeBERTa_Sentiment": {
"悲しみ": 0.12533913552761078… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/moe-speech-plus.moe-speech-20speakers-ljspeech
MoeSpeech 20 Speakers - LJSpeech Format
moe-speech-plus-ljspeechから発話数上位20話者を抽出したデータセットです。
Piper TTSなどのTTSモデル学習に最適化されています。
Dataset Summary
項目
値
話者数
20
総発話数
60,233
フォーマット
LJSpeech
サンプルレート
22050 Hz
言語
日本語
総音声時間
約90.6時間
Speaker Statistics
Speaker ID
Utterances
Internal ID
940de876
4,675
0
2cf01874
4,632
1
bbd90363
3,747
2
1a5a3db8
3,295
3
4e2f4ba6
3,268
4
cc948b89
3,043
5
ad28b91b
3,033
6
ee093a4f
2,957
7… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/moe-speech-20speakers-ljspeech.moe-dict
MOE Taiwanese Dictionary (台語字典)
Word/phrase-level dictionary entry recordings, likely sourced from the Ministry of
Education's official Taiwanese dictionary (臺灣台語常用詞辭典) -- each entry is a single
word or short phrase (not a full sentence), with Han-lo text, a Mandarin gloss, and Tai-lo
romanization.
Dataset Structure
id: numeric dictionary-entry identifier (e.g. 45549), already globally unique.
audio: audio clip (16kHz, mono, 24-bit PCM WAV (majority)).
text:… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/moe-dict.moe-speech-5speakers-ljspeech
moe-speech-5speakers-ljspeech
Japanese multi-speaker TTS dataset in LJSpeech format.
Overview
Speakers: 5
Total utterances: 19,617
Format: LJSpeech (wavs/ + metadata.csv)
Sample rate: 22050 Hz
Language: Japanese
Speaker Statistics
Speaker ID
Utterances
940de876
4,675
2cf01874
4,632
bbd90363
3,747
1a5a3db8
3,295
4e2f4ba6
3,268
Files
metadata.csv - Full metadata (id|speaker|text)
metadata_{speaker_id}.csv - Per-speaker… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/moe-speech-5speakers-ljspeech.tat_moe_cm
TEST
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
nan_tw
Taigi
12.89
12,578
151,234
3.69
3.26
0
0
Total
-
12.89
12,578
151,234
3.69
3.26
0
0
EVAL
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
nan_tw
Taigi
11.68
10,766
134,941
3.91
3.21
0
0
Total
-
11.68
10,766
134,941
3.91
3.21
0
0
TRAIN
Subset
lang_name
hours
n_utts… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/tat_moe_cm.hakkadict_moe_example
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_dp
Hakka_Dapu
0.00
0
0
0.00
0.00
14,455
282,665
hak_hl
Hakka_Hailu
0.00
0
0
0.00
0.00
14,996
294,876
hak_nsx
Hakka_NanSixian
0.00
0
0
0.00
0.00
14,776
291,321
hak_rp
Hakka_Raoping
0.00
0
0
0.00
0.00
14,881
293,945
hak_sx
Hakka_Sixian
0.00
0
0
0.00
0.00
15,010
299,058
hak_za
Hakka_Zhaoan
0.00
0
0
0.00
0.00
12,530
246,820
Total
-
0.00
0
0
0.00
0.00
86,648
1… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hakkadict_moe_example.tat_moe_mc
TEST
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
nan_tw
Taigi
12.86
12,582
151,247
3.68
3.27
0
0
Total
-
12.86
12,582
151,247
3.68
3.27
0
0
EVAL
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
nan_tw
Taigi
11.74
10,769
134,953
3.92
3.19
0
0
Total
-
11.74
10,769
134,953
3.92
3.19
0
0
TRAIN
Subset
lang_name
hours
n_utts… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/tat_moe_mc.hakkadict_moe_word
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_dp
Hakka_Dapu
7.05
11,603
25,919
2.19
1.02
0
0
hak_hl
Hakka_Hailu
4.56
11,403
25,468
1.44
1.55
0
0
hak_nsx
Hakka_NanSixian
4.81
11,676
26,283
1.48
1.52
0
0
hak_rp
Hakka_Raoping
3.68
11,391
25,495
1.16
1.93
0
0
hak_sx
Hakka_Sixian
3.86
11,393
25,519
1.22
1.84
0
0
hak_za
Hakka_Zhaoan
3.31
9,946
21,265
1.20
1.78
0
0
Total
-
27.27
67,412
149,949
1.46
1.53
0
0… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hakkadict_moe_word.moe-speech-plus-ljspeech
MoeSpeechPlus-ljspeech
MoeSpeechPlusのデータセットのデータフォーマットをljspechに沿って変換したもの
huggingface-cli download --repo-type dataset ayousanz/moe-speech-plus-ljspeech --local-dir moe-speech-plus-ljspeech
MoeSpeech
日本語はこちら
このデータセットは、著作権法第三十条の四の情報解析(機械学習等)の目的でのみ使用が許可されています。それ以外の用途での使用はライセンスにより禁止されています。
This dataset is only permitted for use under Article 30-4 of the Copyright Law of Japan for data analysis (such as machine learning) purposes. Any use for purposes other than those… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/moe-speech-plus-ljspeech.
