datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nota
Dataset Card for Nota
Dataset Summary
This data was created by the public institution Nota, which is part of the Danish Ministry of Culture. Nota has a library audiobooks and audiomagazines for people with reading or sight disabilities. Nota also produces a number of audiobooks and audiomagazines themselves.
The dataset consists of audio and associated transcriptions from Nota's audiomagazines "Inspiration" and "Radio/TV". All files related to one reading of one edition… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nota.ivirits-audio-v2-30s
ivrit.ai audio-v2 — 2–30 s segments
ivrit-ai/audio-v2 (>20k hours of Hebrew
audio) cut into 2–30 second speech segments with machine transcripts, ready for ASR
fine-tuning.
How it was built
VAD — Silero VAD (ONNX) over each episode decoded to 16 kHz mono. Speech regions
longer than 30 s are split at the quietest sufficiently-long pause inside the window,
so cuts land in silence rather than mid-word. Regions shorter than 2 s are dropped.
Transcription —… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/ivirits-audio-v2-30s.librispeech_asr_enriched
Dataset Card for librispeech_asr_enriched
synthetic-multilingual-speaker-diarization
Synthetic Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 3417 files
├── all_samples_combined.csv # Complete dataset annotations (with silence)
└── all_visible_combined.csv # Visible dataset annotations (without silence)
Statistics
Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.afvoices-notag
AfVoices Top-20 Speakers without Tags
RobotsMali/afvoices-notag is a small experimental TTS-oriented selection derived from RobotsMali/afvoices, the African Next Voices Bambara speech corpus. It contains the 20 participants with the highest utterance counts and excludes transcripts containing semantic/acoustic annotation tags.
This is the dataset used for RobotsMali's first Bambara VITS experiments. It is not a high-quality studio TTS corpus: the source is spontaneous speech… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices-notag.Knesset-VOX-IPA
Knesset VOX IPA
Hebrew speech dataset derived from Knesset (Israeli Parliament) plenary sessions, enriched with IPA (International Phonetic Alphabet) phoneme transcriptions. Inspired by the methodology of arxiv:2603.01270.
Dataset Description
Long-form Knesset recordings were split into chunks of up to 15 seconds. Each chunk was transcribed to Hebrew text and then processed for IPA phoneme extraction from audio.
Each sample pairs a WAV audio chunk with:
The original… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/Knesset-VOX-IPA.notebooklm_rus
NotebookLM Russian Podcast Dataset
Датасет содержит записи подкастов, сгенерированных с помощью Google NotebookLM на русском языке.
Описание
Голоса: 2 диктора — мужской и женский
Общая длительность: 77 ч 23 мин 22 сек
Количество эпизодов: 417
Формат аудио: WAV, 24 kHz, моно
Язык: русский
Структура датасета
Поле
Тип
Описание
audio
Audio
Аудиозапись эпизода (24 kHz, моно)
transcription
string
Полная текстовая расшифровка эпизода
segments
string… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/notebooklm_rus.Noisy-Voice-Notes
Noisy Voice Notes
A small, hand-curated set of real-world voice notes recorded over roughly a year by a single speaker (me — Daniel Rosehill), captured in everyday environments rather than a studio. Most clips contain meaningful background noise: traffic, café chatter, kitchens, public transport, wind, kids, etc. They are deliberately not clean.
The dataset is intended as a small evaluation / probing set, not a training corpus. Each clip ships with the original transcript plus a… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Noisy-Voice-Notes.dia-Notsofar-test
NOTSOFAR-1 — eval splits (mirror)
Mirror byte-exact des 3 splits du dossier benchmark-datasets/eval_set/ du
repo upstream microsoft/NOTSOFAR (Vinnikov et al. 2024,
CHiME-8 NOTSOFAR-1 Challenge).
Split
Files
Size
GT
240629.1_eval_small
1 360
15.4 GB
— (no GT)
240629.1_eval_small_with_GT
1 899
20.0 GB
✓
240825.1_eval_full_with_GT
4 509
49.2 GB
✓
Total : 7 768 files, ~84.5 GB.
Structure (préservée à l'identique)… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-Notsofar-test.finetuned-hindi-punjabi-denoised
Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 627 files
├── csv/ # Individual CSV annotations - 627 files
├── rttm/ # RTTM format files for diarization - 627 files
├── all_samples_combined.csv # Complete dataset annotations
└── all_samples_combined.rttm # Complete RTTM… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/finetuned-hindi-punjabi-denoised.notaNota lyd- og tekstdata
Datasættet indeholder både tekst- og taledata fra udvalgte dele af Nota's lydbogsbiblotek. Datasættet består af
over 500 timers oplæsninger og medfølgende transkriptioner på dansk. Al lyddata er i .wav-format, mens tekstdata
er i .txt-format.
I data indgår indlæsninger af Notas eget blad "Inspiration" og "Radio/TV", som er udgivet i perioden 2007 til 2022.
Nota krediteres for arbejdet med at strukturere data, således at tekst og lyd stemmer overens.
Nota er en institution under Kulturministeriet, der gør trykte tekster tilgængelige i digitale formater til personer
med synshandicap og læsevanskeligheder, fx via produktion af lydbøger og oplæsning af aviser, magasiner, mv.indonesian-voice-note-dataset
Indonesian Voice Note Dataset
Dataset plan for Indonesian WhatsApp-style voice note intelligence.
Intended Fields
{
"audio_id": "",
"audio_path": "",
"language": "id",
"accent_region": "",
"transcript": "",
"summary": "",
"intent": "",
"urgency": "low|medium|high",
"action_items": [],
"deadline_mentions": [],
"reply_style": "",
"consent": true
}
Collection Rules
Use only consented audio.
Do not include private third-party… See the full description on the dataset page: https://huggingface.co/datasets/terancammuda/indonesian-voice-note-dataset.
