datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pangloss
Dataset Card for [Needs More Information]
Dataset Summary
Two audio corpora of minority languages of China (Japhug and Na), with transcriptions, proposed as reference data sets for experiments in Natural Language Processing. The data, collected and transcribed in the course of immersion fieldwork, amount to a total of about 1,900 minutes in Japhug and 200 minutes in Na. By making them available in an easily accessible and usable form, we hope to facilitate the development… See the full description on the dataset page: https://huggingface.co/datasets/Lacito/pangloss.ds007808-sub01-speechopen-pangolin-preprocessed
ds007808 sub-01 / speechopen / pangolin — preprocessed EEG↔speech windows
Ready-to-train EEG↔speech windows for replicating the scaling experiment of
Sato et al. 2024, "Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours
of EEG Data" (arXiv:2407.07595), built from the public
ds007808 dataset (arXiv:2606.01264).
Slice = subject sub-01, task speechopen (overt speech), device pangolin (128-ch
g.Pangolin) — the rig matching the 175 h paper. Each example is one… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed.teochew_wild
Teochew-Wild:首个正字标注的野外潮州话数据集
本数据集(Teochew-Wild)是从网络上发音清晰、噪声较少的音视频内容中获取的,原始音视频的数据来源为:民生新闻、潮汕讲古、地方电视节目、故事书、抖音自媒体口播等,我借鉴了Emilla提出的数据集自动处理流水线,对原始数据进行归一化、降噪和剪切(部分自动剪切效果差的使用手工修正);
Teochew-Wild总共包括20个发音标准、念错率低的潮汕母语说话人、共12500条音频片段,包含潮州市区、汕头市区、澄海、榕江音、潮安南部等多个区域的口音,语料内容覆盖书面用语与口头用语,并同时提供正字和拼音标注,是首个公开可用、标注准确率高的潮州话数据集,主要面向语音识别和语音合成任务。
文件说明 (File Structure Explanation)
├── label_for_qwen_asr/ # 预处理标签文件夹,完全适配Qwen-ASR模型读取格式
├── README.md # 项目说明文档(本文档)… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_wild.panta_instruct_multi_modal_v1
Panta Instruct Multi-Modal v1
Dataset d'instructions multimodal en français : chaque exemple associe une question
(texte + parole + pictogrammes) à une réponse (texte + pictogrammes).
Colonnes
Colonne
Type
Description
audio
Audio (24 kHz, mono)
Enregistrement de la question (text_input)
text_input
string
Question / instruction
text_output
string
Réponse
pictos_input
list[string]
Identifiants des pictogrammes de la question
pictos_output… See the full description on the dataset page: https://huggingface.co/datasets/audibeal74/panta_instruct_multi_modal_v1.khmer-english-codeswitch-tts
Khmer–English Code-Switch Synthetic Speech
8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech,
generated with VoxCPM2 from code-switch text
manufactured by confirmed lexical substitution over a Khmer–English parallel corpus.
⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a
TTS model, and the code-switch sentences were manufactured by word substitution — they are not
transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.bedvibe-emotional-speechBedVibe Emotional Speech Dataset
Studio-quality emotional speech dataset
• Preview samples available on Hugging Face
• Full commercial dataset available via BedVibe Studio
• Languages currently available: English, Greek
• Additional languages can be recorded on request
• 6 emotions
• 48 kHz / 32-bit float audio
• Professionally recorded in studio conditions
Studio-recorded emotional speech datasets designed for training text-to-speech systems.
Currently available on our website: more than 108… See the full description on the dataset page: https://huggingface.co/datasets/pan82/bedvibe-emotional-speech.pangal
Boli Pangal Data Transcription
About
The Dataset
The current data preview of Pangal is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -
Speech Recordings of 200 sentences in the language.
Transcriptions in IPA and Bangali, English
Translations in English (which also act as prompts for the translation sentences)
Detailed speaker metadata, including their demographic… See the full description on the dataset page: https://huggingface.co/datasets/project-boli/pangal.
