datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hebrew_speech_campus
Data Description
Hebrew Speech Recognition dataset from Campus IL.
Data was scraped from the Campus website, which contains video lectures from various courses in Hebrew.Then subtitles were extracted from the videos and aligned with the audio.Subtitles that are not on Hebrew were removed (WIP: need to remove non-Hebrew audio as well, e.g. using simple classifier).Samples with duration less than 3 second were removed.Total duration of the dataset is 152 hours.Outliers in terms… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/hebrew_speech_campus.DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translateDisclaimer: The original dataset can be found here.
It is published by Digital Divide Data Cambodia (DDD-Cambodia).
License:
Khmer ASR Cultural Dataset's license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0).
Please attribute Digital Divide Data if you use this dataset in any way.
Objective of this dataset
Add English translation: a new column en_translate is added to the original dataset (only from parquet 000 to 159 of the original… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translate.CAMEO
CAMEO: Collection of Multilingual Emotional Speech Corpora
Dataset Description
CAMEO is a curated collection of multilingual emotional speech datasets.
It includes 13 distinct datasets with transcriptions, encompassing a total of 41,265 audio samples.
The collection features audio in eight languages: Bengali, English, French, German, Italian, Polish, Russian, and Spanish.
Example Usage
The dataset can be loaded and processed using the datasets library:
from… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/CAMEO.CAMEO
CAMEO: Collection of Multilingual Emotional Speech Corpora
Dataset Description
CAMEO is a curated collection of multilingual emotional speech datasets.
It includes 13 distinct datasets with transcriptions, encompassing a total of 41,265 audio samples.
The collection features audio in eight languages: Bengali, English, French, German, Italian, Polish, Russian, and Spanish.
Example Usage
The dataset can be loaded and processed using the datasets library:
from… See the full description on the dataset page: https://huggingface.co/datasets/yiyandeng/CAMEO.camoes_SI
camoes_SI
Dataset Description
The camoes_SI dataset is a curated combination of two European
Portuguese sociolinguistic corpora --- Fala Bracarense and
Português Fundamental --- merged into a unified test-only
dataset for evaluating Automatic Speech Recognition (ASR) systems.
All audio is provided as 16 kHz PCM waveforms, accompanied by
speaker metadata and reference transcripts.
This dataset corresponds to the Sociolinguistic Interviews (SI)
category of the CAMÕES… See the full description on the dataset page: https://huggingface.co/datasets/inesc-id/camoes_SI.
