datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NOTSOFAR
Introduction
Welcome to the "NOTSOFAR-1: Distant Meeting Transcription with a Single Device" Challenge.
This repo contains the baseline system code for the NOTSOFAR-1 Challenge.
For more information about NOTSOFAR, visit CHiME's official challenge website
Register to participate.
Baseline system description.
Contact us: join the chime-8-notsofar channel on the CHiME Slack, or open a GitHub issue.
📊 Baseline Results on NOTSOFAR dev-set-1
Values are presented in… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/NOTSOFAR.nota
Dataset Card for Nota
Dataset Summary
This data was created by the public institution Nota, which is part of the Danish Ministry of Culture. Nota has a library audiobooks and audiomagazines for people with reading or sight disabilities. Nota also produces a number of audiobooks and audiomagazines themselves.
The dataset consists of audio and associated transcriptions from Nota's audiomagazines "Inspiration" and "Radio/TV". All files related to one reading of one edition… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nota.ivirits-audio-v2-30s
ivrit.ai audio-v2 — 2–30 s segments
ivrit-ai/audio-v2 (>20k hours of Hebrew
audio) cut into 2–30 second speech segments with machine transcripts, ready for ASR
fine-tuning.
How it was built
VAD — Silero VAD (ONNX) over each episode decoded to 16 kHz mono. Speech regions
longer than 30 s are split at the quietest sufficiently-long pause inside the window,
so cuts land in silence rather than mid-word. Regions shorter than 2 s are dropped.
Transcription —… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/ivirits-audio-v2-30s.not-wake-words-speech-en
not-wake-words-speech-en
Negative (non-wake-word) speech clips, used to measure false accepts for OVOS
wake-word plugins.
Derived from the Multilingual Spoken Words Corpus
(MLCommons), which is built from Mozilla Common Voice and licensed CC-BY-4.0.
This derivative keeps the same licence and attribution requirement.
Produced with support from the NGI0 Commons Fund.
Voice-Note-Audio
Voice Notes Dataset
Dataset Description
This dataset contains real-world voice recordings with transcripts and comprehensive annotations.
Dataset Statistics
Total Entries: 2
Audio Files: 2
Uncorrected Transcripts: 2
Ground Truth Transcripts: 0
Annotation Files: 2
Export Date: 2025-10-27
Dataset Structure
audio/ # Audio recordings (MP3, etc.)
├── 1.mp3
├── 2.mp3
└── ...
transcripts/
├── uncorrected/ # Original STT… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Voice-Note-Audio.MMAU-mini-do-not-useWARNING: The original dataset is revised and pleased refer to new data source. Please refer to: MMAU-v05.15.25: https://github.com/Sakshi113/MMAU
@misc{sakshi2024mmaumassivemultitaskaudio,
title={MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark},
author={S Sakshi and Utkarsh Tyagi and Sonal Kumar and Ashish Seth and Ramaneswaran Selvakumar and Oriol Nieto and Ramani Duraiswami and Sreyan Ghosh and Dinesh Manocha},
year={2024},
eprint={2410.19168}… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/MMAU-mini-do-not-use.guitar-fretboard-notes
Guitar Single-Note Recordings
A dataset of 390 single-note guitar recordings spanning 6 strings and frets 0-12, recorded by two players on acoustic and electric guitars.
Dataset Summary
This dataset contains isolated single-note recordings from a standard-tuned guitar. Each recording captures one note played on a specific string and fret combination, covering the first 12 frets across all 6 strings (78 unique notes per source). The recordings are raw, unprocessed 44100 Hz… See the full description on the dataset page: https://huggingface.co/datasets/collegefishiesd/guitar-fretboard-notes.lofiHipHop
Lofi Dataset 🎵
A bunch of lofi hip hop audio for machine learning purposes. Royalty free.
adapted from jacksonkstenger/lofiHipHop
synthetic-multilingual-speaker-diarization
Synthetic Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 3417 files
├── all_samples_combined.csv # Complete dataset annotations (with silence)
└── all_visible_combined.csv # Visible dataset annotations (without silence)
Statistics
Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.afvoices-notag
AfVoices Top-20 Speakers without Tags
RobotsMali/afvoices-notag is a small experimental TTS-oriented selection derived from RobotsMali/afvoices, the African Next Voices Bambara speech corpus. It contains the 20 participants with the highest utterance counts and excludes transcripts containing semantic/acoustic annotation tags.
This is the dataset used for RobotsMali's first Bambara VITS experiments. It is not a high-quality studio TTS corpus: the source is spontaneous speech… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices-notag.Notcrowy
Audio Dataset Statistics
Overview
Metric
Value
Total audio files
556,667
Total duration
1,024.71 hours (3,688,949 seconds)
Average duration
6.63 seconds
Shortest clip
0.41 seconds
Longest clip
44.97 seconds
Speaker Breakdown
Top 10 Speakers by Clip Count
Speaker
Clips
Duration
% of Total
Despina
60,150
118.07 hours
11.5%
Sulafat
31,593
58.15 hours
5.7%
Achernar29,889
54.53 hours
5.3%
Autonoe
27,897… See the full description on the dataset page: https://huggingface.co/datasets/mobinx/Notcrowy.notebooklm_rus
NotebookLM Russian Podcast Dataset
Датасет содержит записи подкастов, сгенерированных с помощью Google NotebookLM на русском языке.
Описание
Голоса: 2 диктора — мужской и женский
Общая длительность: 77 ч 23 мин 22 сек
Количество эпизодов: 417
Формат аудио: WAV, 24 kHz, моно
Язык: русский
Структура датасета
Поле
Тип
Описание
audio
Audio
Аудиозапись эпизода (24 kHz, моно)
transcription
string
Полная текстовая расшифровка эпизода
segments
string… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/notebooklm_rus.minecraft-note-block
Minecraft Note Block
Literally just note block samples.
SententicDataTTS
SententicDataTTS
A Hebrew and English TTS dataset with male and female speakers, resampled to 44.1kHz and time-stretched (slowed).
Audio Generation
slow_44K.7z — generated using Chatterbox
Mamre_generated.7z — generated using MamreTTS
Contents
slow_44K.7z
Audio files resampled to 44.1kHz and time-stretched (slowed), containing:
data/ — audio files (WAV, 44.1kHz, slowed)
CSV metadata files per speaker:… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/SententicDataTTS.vaani-hindibelt-notranscript-concattrain_de_en_de_not_similar_final
Dataset Card for "train_de_en_de_not_similar_final"
More Information needed
Noto_Mamikoassetstrain_en_de_en_not_similar_final
Dataset Card for "train_en_de_en_not_similar_final"
More Information needed
dia-Notsofar-test
NOTSOFAR-1 — eval splits (mirror)
Mirror byte-exact des 3 splits du dossier benchmark-datasets/eval_set/ du
repo upstream microsoft/NOTSOFAR (Vinnikov et al. 2024,
CHiME-8 NOTSOFAR-1 Challenge).
Split
Files
Size
GT
240629.1_eval_small
1 360
15.4 GB
— (no GT)
240629.1_eval_small_with_GT
1 899
20.0 GB
✓
240825.1_eval_full_with_GT
4 509
49.2 GB
✓
Total : 7 768 files, ~84.5 GB.
Structure (préservée à l'identique)… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-Notsofar-test.test_en_de_en_not_similar_final
Dataset Card for "test_en_de_en_not_similar_final"
More Information needed
train_de_en_de_not_similar_injected_1st_1000
Dataset Card for "train_de_en_de_not_similar_injected_1st_1000"
More Information needed
arazn_codeSwitched_mp3_full_notLowertrain_en_de_en_not_similar_okraje_1st_1000
Dataset Card for "train_en_de_en_not_similar_okraje_1st_1000"
More Information needed
upto_2025_may_15_stc_0_38068_notNormalized_saskia_only_datasetMoisesup_testfinetuned-hindi-punjabi-denoised
Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 627 files
├── csv/ # Individual CSV annotations - 627 files
├── rttm/ # RTTM format files for diarization - 627 files
├── all_samples_combined.csv # Complete dataset annotations
└── all_samples_combined.rttm # Complete RTTM… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/finetuned-hindi-punjabi-denoised.test_de_en_de_not_similar_final
Dataset Card for "test_de_en_de_not_similar_final"
More Information needed
ATCO2-ASR
