datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epstractor-raw
Epstractor: Epstein Archives Dataset
A comprehensive archive of documents, images, audio, and video files from multiple Epstein-related releases, including estate records and Department of Justice materials obtained through FOIA requests.
Dataset Description
This dataset contains 59,420 files totaling 115.23 GB from three major document releases, plus 2 large videos (40GB) available via a separate config:
Epstein Estate 2025-09: 5 files, 0.09 GB
Epstein Estate 2025-11:… See the full description on the dataset page: https://huggingface.co/datasets/public-records-research/epstractor-raw.avian-dawn-chorus-recordingsRehosting of the Zenodo dataset: Audio tagging of avian dawn chorus recordings in California, Oregon, and Washington dataset, Weldy et al (2024)
recorded_evaluation_dataset2lumi-record-smh
lumi-record-smh
Vietnamese smart-home voice recording dataset exported from lumi-record-voice-SMH.
Generated at: 2026-07-29T11:45:33
Configs
mien-bac: Northern voices.
mien-trung: Central voices.
mien-nam: Southern voices.
All configs expose one split: test.
Columns
The dataset intentionally exposes only the columns needed for speech, speaker, and recording-context analysis.
Column
Type
Description
audio
Audio
Embedded WAV audio. In the… See the full description on the dataset page: https://huggingface.co/datasets/luvox-ai/lumi-record-smh.audio-record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 2592,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"audio_files_size_in_mb": 100,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AdamMalyshev/audio-record-test.Arabic_audio_recordingsCan test it.
https://huggingface.co/datasets/a1anas1a/Arabic_audio_recordings/edit/main/README.md
https://fastminify.com/en/minify-js
4tts-recordingschinese_speech_self-recordedrecord_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 433,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"audio_files_size_in_mb": 100,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/record_test.kws-recordings-soundailabel
kws-recordings-soundailabel
This repository serves as the KWS audio dataset repo for label_kws training.
Layout
audio/
manifests/
docs/
README.md
dataset_summary.json
Source of Truth
manifests/full_manifest.csv
manifests/kws_multiclass_manifest.csv
manifests/binary_emergency_manifest.csv
manifests/review_queue.csv
For compatibility, the same derived files are also available under data/derived/.
Current Summary
rows with local audio present: 796… See the full description on the dataset page: https://huggingface.co/datasets/dusen0528/kws-recordings-soundailabel.call_recordings-v1-818test-recordingsprocessed_phone_recordingsparallel-recordings
Parallel Recordings
This dataset is a small corpus of audio recorded in parallel.
This data can be used to quantify the difference in recording quality between different audio devices.
An example of this quantification using WER is shown below.
Organization
There are two separate experiments.
Personal Microphone
Speakerphone
The original data for both experiments are located in thefull_length_audio directory.
Each experiment was performed by recording simultaneously on… See the full description on the dataset page: https://huggingface.co/datasets/lescidium/parallel-recordings.audio_recordings
Audio Recordings
These are some personal and open source audio recordings. Feel free to use.
voice-recorder-test3-demoaudio_recordSA-zenande-recordsrecording-deletevoice_model_recordingLibriReplay-DOA
LibriReplay-DOA (Anonymous Submission)
Overview
LibriReplay-DOA is a multi-channel multi-speaker replay dataset designed for evaluating
robust speech processing systems under realistic room playback conditions.
The dataset contains replay recordings captured in real rooms under
multiple playback configurations (DOA settings). Each session includes
multiple overlapping speakers.
This dataset is released for peer-review purposes.
Dataset Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/real-recordings/LibriReplay-DOA.video_recordingTheresa-Recordingwake-word-recordingsMay2024-mixedrecorded_audioRecordWritingRecording_reports_on_work_done
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/KrugDen/Recording_reports_on_work_done.solo-clarinet-classical-recordings
Solo Clarinet AI Training Dataset – Isolated Monophonic Recordings (Sample)
High-quality isolated clarinet recordings for AI audio modeling and research.
This dataset is part of a larger collection of solo clarinet recordings designed for AI and audio machine learning applications.
It provides clean, high-quality monophonic recordings of a real acoustic instrument, suitable for tasks such as timbre modeling, music generation, and sound synthesis.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/clarinet-ai-datasets/solo-clarinet-classical-recordings.nnh-record-queueATC-recordings
