datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VocalVerse-datasetElaina_WanderingWitch_audio_JA
伊蕾娜 语音数据集 (Elaina Voice Audio Dataset)
角色介绍
伊蕾娜(Elaina / イレイナ) 是来自 《魔女之旅》(Majo no Tabitabi / 魔女の旅々) 的主角。
她是一位银发琉璃瞳的旅人魔女,带着温柔的微笑造访各地,见证人间百态。性格冷静又带点俏皮,声音柔和动听,由 本渡楓(Kaede Hondo) 配音。
"这位戴着魔女证明的胸针,飘扬着灰色秀发,那美貌与才能的光辉,让太阳都不禁眯起双眼的美女,到底是谁呢?没错,就是我。"— 伊蕾娜
数据集概述
本数据集收录了 伊蕾娜 日语配音 的音频切片及对应文本。
基本信息
项目
内容
角色
伊蕾娜 (Elaina)
作品
魔女之旅 (Majo no Tabitabi)
配音演员
本渡楓 (Kaede Hondo)
配音语言
日语(JA)
音频来源
B站 / Bilibili
数据格式
Parquet + MP3/WAV
数据预览… See the full description on the dataset page: https://huggingface.co/datasets/yeeko/Elaina_WanderingWitch_audio_JA.comfyui-wan22-assetsthai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/thai-dialect-isan-dataset.enwaucymraegThe training and development set sentences are taken from CoVoST and have been compared to all validated sentences in the Welsh Common Voice data to ensure none of the already recorded sentences will be used here. Then all sentences containing personal names have been extracted and replaced with a randomly generated name using the Faker library and a custom Welsh names list. The sentences were then recorded by 26 volunteers from North-West Wales, 15 women, 10 men and one non-binary person.… See the full description on the dataset page: https://huggingface.co/datasets/wanasash/enwaucymraeg.thai-ser
🇹🇭 THAI-SER Dataset 🎭
[📝 Paper (preprint)]
Published by: AI Research Institute of Thailand (AIResearch)
In collaboration with:
Vidyasirimedhi Institute of Science and Technology (VISTEC)
Digital Economy Promotion Agency (depa)
Department of Computer Engineering, Faculty of Engineering, Chulalongkorn University
Department of Dramatic Arts, Faculty of Arts, Chulalongkorn University
Sponsored by: Advanced Info Services Public Company Limited (AIS), and Siam Commercial… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/thai-ser.MuChin1khttps://github.com/CarlWangChina/MuChin
whisper-large-v2-eval-cvModel: openai/whisper-large-v2
Test Set: DewiBrynJones/commonvoice_18_0_cy
Split: test
WER: 41.304748
CER: 16.138398
WanJuanSiLu-Multimodal-5Languages
WanJuan·SiLu Multimodal Multilingual Corpus
🌏Dataset Introduction
The newly upgraded "Wanjuan·Silk Road Multimodal Corpus" brings the following three core improvements:
The number of languages has been significantly expanded: Based on the five open-source languages of "Wanjuan·Silk Road", namely Arabic, Russian, Korean, Vietnamese, and Thai, "Wanjuan·Silk Road Multimodal" has added three scarce corpus data of Serbian, Hungarian, and Czech, and uses the above… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuanSiLu-Multimodal-5Languages.SaMoyeSVCwhisper-large-v3-ec-eval-cvModel: wanasash/whisper-large-v3-ec
Test Set: DewiBrynJones/commonvoice_18_0_cy
Split: test
WER: 38.786166
CER: 11.659708
corpus-siarad-test-setWang_Leehom_Music_Class_audio_sampleVaani-assamese-wancho-nepali-lg-English-no-transcript0fleurs_demowanglihong-matchedwhisper-large-v2-ec-eval-cvModel: wanasash/whisper-large-v2-ec
Test Set: DewiBrynJones/commonvoice_18_0_cy
Split: test
WER: 38.194402
CER: 11.612194
Vaani-wancho-nepali-majority-lg-English-with-transcriptwhisper-large-v3-ec-eval-ecModel: wanasash/whisper-large-v3-ec
Test Set: wanasash/enwaucymraeg
Split: test
WER: 28.331177
CER: 7.949406
11537436_WangYiLinLinda
Dataset Description
This dataset contains 3 hours of clear and high-quality Mandarin audio at a 44.1kHz sampling rate,sourced from a language laboratory, along with precise textual annotations. The recording was conducted in a manner of 10 sentences per group. Subsequently, the audio was edited and selected to ensure that the volume of each sentence is between 0.3 and 0.7 and that the silent intervals before and after each sentence are between 100 and 200 milliseconds, in order to… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/11537436_WangYiLinLinda.nepali_asr_evaluation_data081000n-demo-ultimatesichuanaudiomc
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
Audio MultiChallenge is an open-source benchmark to evaluate E2E spoken dialogue systems under natural multi-turn interaction patterns. Building on the text-based MultiChallenge framework, which evaluates Inference Memory, Instruction Retention, and Self Coherence, we introduce a new axis Voice Editing that tests robustness to mid-utterance speech repairs and backtracking. We… See the full description on the dataset page: https://huggingface.co/datasets/wangyueyiiiiiii/audiomc.SASLM-demo-audiowhisper-large-v2-ec-eval-ecModel: wanasash/whisper-large-v2-ec
Test Set: wanasash/enwaucymraeg
Split: test
WER: 27.899957
CER: 8.233039
whisper-large-v2-eval-ecModel: openai/whisper-large-v2
Test Set: wanasash/enwaucymraeg
Split: test
WER: 46.959897
CER: 17.439632
SparkTTS_Wang_Leehom_Ad_wavMuChin-v2-6066https://github.com/CarlWangChina/MuChin-V2-6066
