datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ashraq-esc50-1-dog-example
Dataset Card for "ashraq-esc50-1-dog-example"
More Information needed
Taiwanese-Minnan-Example-Sentences
Taiwanese Minnan Example Sentences
The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems.
Dataset Features
Source: Ministry of Education, Taiwan (Sutian Resource Center)
Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.voice-focus-examples
Voice Focus Examples
Collection of examples for accurate foreground speaker transcription.
Details
Curated by: Joschka Wohlgemuth
Funded by: ai-coustics GmbH
Contact:
Web: https://ai-coustics.com
Omni-DuplexEval-ExamplesOmni-DuplexEval-Examples is a curated subset of Omni-DuplexEval for qualitative visualization and paper demonstration. Each task contains 5 representative samples, covering all benchmark scenarios and task types.
This subset is intended for illustrative purposes in the paper and supplementary materials. The annotation format and data structure are consistent with the full benchmark.
Omni-DuplexEval
Omni-DuplexEval is a benchmark for evaluating real-time duplex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/foragi/Omni-DuplexEval-Examples.multilingual_examplesexample_mmdata_mnbvc
mnbvc mm dataset v2.1
MNBVC 多模态语料数据格式。原链接:https://huggingface.co/datasets/wanng/example_mmdata_mnbvc
参考实现:mm_template_mnbvc
的 mmdata_block.BLOCK_SCHEMA。schema 以那份代码为准,这个数据集是它的示例产物。
字段
字段名称
类型
字段说明
可选
实体ID
string
数据的唯一标识符。用于在数据集中确定是哪一条数据。在单个数据集中确定一条数据的实体对象。
必选
md5
string
内容的 md5,用于去重与完整性校验
必选
块ID
int32
一个实体对象内的标识符。用于确定一条数据内的一个部分数据。parquet 行的最小单元。
必选
块类型
string
用于保存块的类别。类别的含义为「模态」。取值见下
必选
扩展字段
string
用于保存块的元信息。为可以被成功 load 的 json 字符串。后期可继续扩展
必选… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/example_mmdata_mnbvc.commonvoice22-sidon-xcodec2-examplestarrail-voice-snac-exampleexample_TTSikema_dictionary_examples_datasetstarrail-voice-xcodec2-exampleimagebind-example-datahakka_elearning_example_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Dapu
客語_大埔
7.50
5,040
69,667
5.36
2.58
Hakka_Hailu
客語_海陸
6.93
5,055
73,544
4.94
2.95
Hakka_Raoping
客語_饒平
4.26
3,202
40,861
4.79
2.66
Hakka_Sixian
客語_四縣
3.43
2,759
36,159
4.48
2.92
Hakka_Zhaoan
客語_詔安
7.59
5,039
70,183
5.43
2.57
Total
-
29.72
21,095
290,414
5.07
2.71
rapnic-example
RAPNIC Dataset (example)
Dataset Description
This is an example of the full dataset, yet to be published, with 10 audio examples for 72 speakers.
RAPNIC (Reconeixement Automàtic de la Parla No Intel·ligible en Català) is a Catalan speech corpus collected from individuals with speech disorders, specifically cerebral palsy and Down syndrome.
This dataset was collected to develop and improve automatic speech recognition (ASR) systems that are accessible to people with speech… See the full description on the dataset page: https://huggingface.co/datasets/CLiC-UB/rapnic-example.sample_examplelibrispeech-asr-exampleexample_dataset
Dataset Card for "example_dataset"
More Information needed
example-codecexample-dialogue-qahinglish-audio-turn-end-example
Hinglish Audio Turn End Example
Dataset of hinglish audio snippets generated with SarvamAI bulbul:v3 TTS.
Each example is a customer service utterance with metadata about the speaker
gender, whether the utterance is complete or incomplete (trailoff / connective /
filler ending), and the rendered audio file.
Fields
field
type
description
script
string
hinglish transcript (Devanagari + Latin script mixed)
audio
audio
mp3 speech clip generated by… See the full description on the dataset page: https://huggingface.co/datasets/Shrey160/hinglish-audio-turn-end-example.exampletts-dataset-example
TTS Dataset Example
Demo output of the pipeline at
github.com/ulaspolat/tts-dataset-creation
60 segments, 24 kHz mono. Columns: audio, text.
Source: A Joke by Anton Chekhov, read for LibriVox.
hakkadict_moe_example
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_dp
Hakka_Dapu
0.00
0
0
0.00
0.00
14,455
282,665
hak_hl
Hakka_Hailu
0.00
0
0
0.00
0.00
14,996
294,876
hak_nsx
Hakka_NanSixian
0.00
0
0
0.00
0.00
14,776
291,321
hak_rp
Hakka_Raoping
0.00
0
0
0.00
0.00
14,881
293,945
hak_sx
Hakka_Sixian
0.00
0
0
0.00
0.00
15,010
299,058
hak_za
Hakka_Zhaoan
0.00
0
0
0.00
0.00
12,530
246,820
Total
-
0.00
0
0
0.00
0.00
86,648
1… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hakkadict_moe_example.starrail-voice-ann-examplehindi_dataset_small_exampleexample_group_1example-dialogueemilia-cjk-xcodec2-exampleemilia-yodas-cjk-xcodec2-examplegame-novel-xcodec2-example
