datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ravnursson_asr
Dataset Card for ravnursson_asr
Dataset Summary
The corpus "RAVNURSSON FAROESE SPEECH AND TRANSCRIPTS" (or RAVNURSSON Corpus for short) is a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications in the language that is spoken at the Faroe Islands (Faroese). It was curated at the Reykjavík University (RU) in 2022.
The RAVNURSSON Corpus is an extract of the "Basic Language Resource Kit 1.0" (BLARK 1.0) [1] developed… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/ravnursson_asr.AIShell
Dataset Card for "Aishell1"
More Information needed
Human-Animal-Cartoon
Human-Animal-Cartoon dataset
Our Human-Animal-Cartoon (HAC) dataset consists of seven actions (‘sleeping’, ‘watching tv’, ‘eating’, ‘drinking’, ‘swimming’, ‘running’, and ‘opening door’) performed by humans, animals, and cartoon figures, forming three different domains. We collect 3381 video clips from the internet with around 1000 for each domain and provide three modalities in our dataset: video, audio, and pre-computed optical flow.
The dataset can be used for Multi-modal Domain… See the full description on the dataset page: https://huggingface.co/datasets/hdong51/Human-Animal-Cartoon.carva-audio-libraryCaReCoS
CaReCoS
A medical acoustic question-answering dataset for reasoning over mel spectrograms
of heart, lung, and cough sounds. Each record provides a clinical question, the
mel-spectrogram image of a recording, a ground-truth answer, and the
recording's clinical metadata.
The task is purely visual: a model receives the spectrogram image together with the
question and must reason over the spectrogram to produce the answer. The raw audio is
not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.plug_socket_mixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 15095,
"total_tasks": 1,
"total_videos": 150,
"total_audio": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_mixed.in_car_commands_26
Dataset Card for "in_car_commands_26"
More Information needed
plug_socket_single_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 11536,
"total_tasks": 1,
"total_videos": 90,
"total_audio": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_single_2.plug_socket_movedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 44,
"total_frames": 15772,
"total_tasks": 1,
"total_videos": 132,
"total_audio": 132,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:44"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_moved.VoiceTrace-BenchVoiceTrace-Bench
VoiceTrace is a benchmark and unified framework for who-said-what speech retrieval: given a natural-language query about a speaker's identity or what they said, retrieve the matching audio document. Unlike conventional speaker verification or diarization benchmarks, VoiceTrace evaluates retrieval jointly over who is speaking and what is being said, across both single-speaker and multi-speaker conversational recordings.
This repository hosts the VoiceTrace-Bench… See the full description on the dataset page: https://huggingface.co/datasets/cara-ai/VoiceTrace-Bench.figli-e-napule-mediachm150_asr
Dataset Card for chm150_asr
Dataset Summary
The CHM150 is a corpus of microphone speech of mexican Spanish taken from 75 male speakers and 75 female speakers in a noise environment of a "quiet office" with a total duration of 1.63 hours.
Speakers were encouraged to respond between some pre selected open questions or they could also describe a particular painting showed to them in a computer monitor. By so, the speech is completely spontaneous and one can see it in the… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/chm150_asr.plug_socket_singleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 65,
"total_frames": 24999,
"total_tasks": 1,
"total_videos": 195,
"total_audio": 195,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:65"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_single.voice-of-care-health-dataset
Voice of Care AI for Global Health Benchmark Dataset
Overview
This dataset contains spoken Hausa Health datasets with rich annotations covering emotion, intent, speaker demographics, and dialect variation, intended for speech and NLP research.
Dataset Summary
Property
Details
Language
Hausa
Modality
Audio + Text
Task(s)
e.g. Speech Recognition, Emotion Detection, Dialect Identification
Version
1.0.0
🛠️ Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Data-Science-Nigeria/voice-of-care-health-dataset.bengali-telecom-customer-care-speech-v2
Bengali Telecom Customer Care Synthetic Speech Dataset v2
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.void-carousel
VOID CAROUSEL
Dark psychedelic trance from graveyard orbit.
Fictional entity
VOID CAROUSEL is a fictional sentient derelict orbital carousel in the Sonic Forage universe: an abandoned amusement machine circling a dead world, translating its failing motors, empty passenger rings and intercepted signals into imagined psychedelic trance. It is not a human performer; its visual identity depicts no human likeness. This is an original fictional characterization, not a… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/void-carousel.dummy_corpus_asr_esThis is an example of a repository where the audio files are not compressed in tar files.
Human-Animal-Cartoon-PC-VAin_car_commands_60
Dataset Card for "in_car_commands_60"
More Information needed
plug_socket_single_slowThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 24939,
"total_tasks": 1,
"total_videos": 150,
"total_audio": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_single_slow.cartesia-sonic-preview-ztts1-zero-shot-sample
Cartesia Sonic on ZTTS1 zero-shot — sample with reference audio
100 utterances per language (700 rows) from the
zero-shot subsets of ZTTS1-Eval, synthesized with Cartesia Sonic (preview) in
voice-cloning mode.
Unlike the full set, every row carries the reference recording as well as the
synthesized clip, so a take can be compared against the voice it was cloning
without checking out the benchmark.
Columns
column
meaning
audio
the clip the model produced… See the full description on the dataset page: https://huggingface.co/datasets/jaeyong2/cartesia-sonic-preview-ztts1-zero-shot-sample.asoul_carol
声音数据
数据来源为asoul的珈乐 22年5月~21年6月的大部分录播时长共5小时 无内容标记已完成响度匹配数据在carol_fast_lzma2.zip里压缩算法是fast lzma2 太旧的解压软件可能不支持字母s开头的音频是歌声数据,量少质量低,建议删除无授权,侵删
2025.2备注
都2025年了还能每月五十个下载,都是神人了💧
plug_socket_single_paddingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 22752,
"total_tasks": 1,
"total_videos": 150,
"total_audio": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_single_padding.car-crash-audio-cc
Car Crash Audio (Creative Commons)
46 clips (up to 40s each, varied length) sourced from
38 YouTube videos licensed Creative Commons -- Attribution,
found via search for "car crash audio".
Attribution
CC BY requires attribution on reuse. Full per-clip attribution (title, channel, source URL) is in
metadata.csv. Summary of unique source videos:
Car Crash Sound Effect | Realistic Impact Audio for Films, Shorts & Game Development -- Stock Media (CC BY)
Car Crash… See the full description on the dataset page: https://huggingface.co/datasets/Titung/car-crash-audio-cc.violence_contextSmartHearingAids-data
Semantic Hearing
This repository provides code for the binaural target sound extraction model proposed in the paper, Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables, presented at UIST'23. This model helps us create systems that let you control what you want to hear in the environment, in real-time, using noise-cancelling earbuds & headphones.
https://github.com/vb000/SemanticHearing/assets/16723254/f1b33d8c-179a-4d50-92aa-6a99dde696d0
Conda environment… See the full description on the dataset page: https://huggingface.co/datasets/carankt/SmartHearingAids-data.carnatic-ragasdimex100_light
Dataset Card for dimex100_light
Dataset Summary
The DIMEx100 LIGHT Corpus (DL) is a reduced version of the DIMEx100 Corpus (D100). DL was created in 2016 by Carlos Daniel Hernández Mena, with the aim of facilitating the use of the DIMEx100 Corpus in various automatic speech recognition systems.
The most important differences between DIMEx100 LIGHT and the original are:
The DL only contains audio files and transcriptions, unlike the D100 which contains pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/dimex100_light.carl_johnson_voice_pack
Carl Johnson Voice Pack/Dataset
record_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 433,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"audio_files_size_in_mb": 100,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/record_test.
