datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pangloss
Dataset Card for [Needs More Information]
Dataset Summary
Two audio corpora of minority languages of China (Japhug and Na), with transcriptions, proposed as reference data sets for experiments in Natural Language Processing. The data, collected and transcribed in the course of immersion fieldwork, amount to a total of about 1,900 minutes in Japhug and 200 minutes in Na. By making them available in an easily accessible and usable form, we hope to facilitate the development… See the full description on the dataset page: https://huggingface.co/datasets/Lacito/pangloss.khmer-tts-processed
Khmer TTS Processed
Preprocessed, Fish Speech-ready Khmer speech data derived from
DDD-Cambodia/khm-asr-cultural
(train split). Built for fine-tuning Fish Speech
(fishaudio/openaudio-s1-mini) for Khmer text-to-speech / voice cloning —
see Panhapich/Tuna-TTS for the
resulting model checkpoints.
Processing pipeline
Raw audio + transcripts were pulled from the source dataset and run through:
Export/merge raw clips into a single manifest
Automated audio-quality (QC)… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-tts-processed.teochew_wild
Teochew-Wild:首个正字标注的野外潮州话数据集
本数据集(Teochew-Wild)是从网络上发音清晰、噪声较少的音视频内容中获取的,原始音视频的数据来源为:民生新闻、潮汕讲古、地方电视节目、故事书、抖音自媒体口播等,我借鉴了Emilla提出的数据集自动处理流水线,对原始数据进行归一化、降噪和剪切(部分自动剪切效果差的使用手工修正);
Teochew-Wild总共包括20个发音标准、念错率低的潮汕母语说话人、共12500条音频片段,包含潮州市区、汕头市区、澄海、榕江音、潮安南部等多个区域的口音,语料内容覆盖书面用语与口头用语,并同时提供正字和拼音标注,是首个公开可用、标注准确率高的潮州话数据集,主要面向语音识别和语音合成任务。
文件说明 (File Structure Explanation)
├── label_for_qwen_asr/ # 预处理标签文件夹,完全适配Qwen-ASR模型读取格式
├── README.md # 项目说明文档(本文档)… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_wild.ds007808-sub01-speechopen-pangolin-preprocessed
ds007808 sub-01 / speechopen / pangolin — preprocessed EEG↔speech windows
Ready-to-train EEG↔speech windows for replicating the scaling experiment of
Sato et al. 2024, "Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours
of EEG Data" (arXiv:2407.07595), built from the public
ds007808 dataset (arXiv:2606.01264).
Slice = subject sub-01, task speechopen (overt speech), device pangolin (128-ch
g.Pangolin) — the rig matching the 175 h paper. Each example is one… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed.khmer-english-codeswitch-tts
Khmer–English Code-Switch Synthetic Speech
8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech,
generated with VoxCPM2 from code-switch text
manufactured by confirmed lexical substitution over a Khmer–English parallel corpus.
⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a
TTS model, and the code-switch sentences were manufactured by word substitution — they are not
transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.panta_instruct_multi_modal_v1
Panta Instruct Multi-Modal v1
Dataset d'instructions multimodal en français : chaque exemple associe une question
(texte + parole + pictogrammes) à une réponse (texte + pictogrammes).
Colonnes
Colonne
Type
Description
audio
Audio (24 kHz, mono)
Enregistrement de la question (text_input)
text_input
string
Question / instruction
text_output
string
Réponse
pictos_input
list[string]
Identifiants des pictogrammes de la question
pictos_output… See the full description on the dataset page: https://huggingface.co/datasets/audibeal74/panta_instruct_multi_modal_v1.khmer-english-codeswitch-tts-llm
Khmer–English Code-Switch Synthetic Speech (LLM-authored)
19,825 utterances / 21.7 hours of synthetic Khmer–English code-switched speech at
16 kHz, generated with VoxCPM2 from code-switch
sentences written by an LLM and validated programmatically.
⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS
model, and every sentence was written by a language model — they are not transcripts of anything
a person said. It is intended as an… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts-llm.bedvibe-emotional-speechBedVibe Emotional Speech Dataset
Studio-quality emotional speech dataset
• Preview samples available on Hugging Face
• Full commercial dataset available via BedVibe Studio
• Languages currently available: English, Greek
• Additional languages can be recorded on request
• 6 emotions
• 48 kHz / 32-bit float audio
• Professionally recorded in studio conditions
Studio-recorded emotional speech datasets designed for training text-to-speech systems.
Currently available on our website: more than 108… See the full description on the dataset page: https://huggingface.co/datasets/pan82/bedvibe-emotional-speech.teochew_gang-gou
潮州讲古数据集(teochew_gang-gou dataset)
Teochew-extLa 数据集的讲古部分。
The "Gang-Gou"(Teochew Traditional Storytelling) part of Teochew-extLa dataset. ("extLa" == "extensible" + "large")
example 1
许昣时 是叫做书斋,了,个名 叫做稻花庄。 (In the past, [it] was called 'Shuzhai' (Book Room/Study Room), then, the name was called 'Daohua Zhuang' (Rice Blossom Cottage).)
example 2
坫觅、走掠,我 唺唺去 坫在许地下室 介 内畔。 (For hide-and-seek, I often hid inside that basement.)
属性 (Attribute)
值 (Value)… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_gang-gou.kgs_ksc_fom
Khmer ASR Phase 1 Cache
This repository contains cached training artifacts used for Khmer ASR backbone evaluation.
Source Datasets
seanghay/khmer_grkpp_speech
seanghay/km-speech-corpus
KrorngAI/fleurs_openslr42_mpwt
Contents
combined_dataset.pt
Usage
import torch
data = torch.load("combined_dataset.pt")
pangal
Boli Pangal Data Transcription
About
The Dataset
The current data preview of Pangal is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -
Speech Recordings of 200 sentences in the language.
Transcriptions in IPA and Bangali, English
Translations in English (which also act as prompts for the translation sentences)
Detailed speaker metadata, including their demographic… See the full description on the dataset page: https://huggingface.co/datasets/project-boli/pangal.
