datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hit-asrpersian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.danish-asr-leaderboard
Open Danish ASR Leaderboard — Results
Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models.
Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.kupe-asr-en-data
kupe-asr-en-mini-150m — data
Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly):
raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this.
mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this.
Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state.
from datasets import load_dataset
ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train")
sanskrit-asr-84danish-asr-verified
danish-asr-verified
ALL rows of syvai/danish-asr-unified transcribed by the
syv-transcribe ensemble (hviske-v5.3 + hviske-v5, confidence-weighted ROVER),
each annotated with:
verified — True when the ensemble independently reproduced the reference
exactly (compared after lowercasing, punctuation-strip, whitespace-collapse).
Two independent witnesses agree => near-certain label.
wer_teacher_vs_ref / cer_teacher_vs_ref — word/character error rate
between normalized teacher output… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-verified.librispeech_asr_sliced
Librispeech Slices
Description
Librispeech is a large corpus of read English utterances derived from the LibriVox public domain audiobook project.
It was assembled to assist in Automatic Speech Recognition tasks, and contains approximately 1000 Hours of utterances recorded at 16kHz.
A subset of the original Librispeech dataset was created to support the development of automatic audio scene creation in the Treble SDK environment.
To better imitate the natural flow of… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/librispeech_asr_sliced.Indonesian-ASR-11-Class-Dataset
Indonesian ASR 11-Class Dataset
Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts.
Dataset summary
Audio files: 104,500 WAV files
Real/human recordings: 104,368
Synthetic repair files: 132
Sentence classes: 11 Indonesian sentence categories
Canonical balanced sentence slots: 209 (11 categories × 19 retained slots)
Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs*
Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.ASR-CYGNSS-HGMM-Reproducibility
ASR CYGNSS-SMAP H-GMM reproducibility repository
This public dataset repository stores derived CYGNSS-SMAP collocations, model
checkpoints, evaluation statistics, and manuscript figures for the manuscript
on year-adaptive CYGNSS sea-surface wind retrieval.
Upstream public data
CYGNSS Level-2 Science Data Record v3.2 (CYGNSS_L2_V3.2), NASA PO.DAAC,
DOI: 10.5067/CYGNS-L2X32.
JPL SMAP Level-2B CAP Sea Surface Salinity and extreme-wind product v5.0… See the full description on the dataset page: https://huggingface.co/datasets/wuff-mann/ASR-CYGNSS-HGMM-Reproducibility.robot-meet-gemma-record-medicineThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 13,
"total_frames": 6119,
"total_tasks": 1,
"total_videos": 26,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:13"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AS-Robotics/robot-meet-gemma-record-medicine.sanskrit-asr-84-evalsalt-asr-data-transcriptionsPick-and-Place-dataset_so100_pickplaceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 3764,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AS-Robotics/Pick-and-Place-dataset_so100_pickplace.librispeech_asr_dummy
Dataset Card for librispeech_asr_dummy
Dataset Summary
This is a truncated version of the LibriSpeech dataset. It contains 20 samples from each of the splits. To view the full dataset, visit: https://huggingface.co/datasets/librispeech_asr
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been… See the full description on the dataset page: https://huggingface.co/datasets/sanchit-gandhi/librispeech_asr_dummy.sd_asr_synthesis_datavoxpopuli_asr_norm_curatorlibrispeech960-encodec1024_asr
Dataset Card for "librispeech960-encodec1024_asr"
More Information needed
indicvoices-v1asd_asr_synthesis_data_v0_less_silencenb-asr-numerics-harvested
Norwegian Bokmål Numeric Expression Harvesting Dataset
This dataset contains cleaned, high-quality Norwegian Bokmål sentences containing numeric expressions harvested from both the Norwegian Colossal Corpus (NbAiLab/NCC) and the FineWeb-2 Norwegian Bokmål subset (HuggingFaceFW/fineweb-2 config nob_Latn).
This is a combined high-volume intermediate dataset built for the first stage of a template-based Norwegian synthetic speech (TTS) generation pipeline to improve number/digit… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-harvested.librispeech960-wavlm-large-km1000_asr
Dataset Card for "librispeech960-wavlm-large-km1000_asr"
More Information needed
condensed_librispeech_asr
Condensed LibriSpeech ASR
This dataset is a condensed version of the LibriSpeech ASR dataset, created by subsampling approximately 10% of the original data from each split. It is intended for quick experimentation, prototyping, and debugging when working with Automatic Speech Recognition (ASR) tasks.
Dataset Details
Original Dataset: LibriSpeech ASR
Condensation Ratio: Approximately 10% of the full dataset
Splits Included:
train.clean.100
train.clean.360
train.other.500… See the full description on the dataset page: https://huggingface.co/datasets/nyalpatel/condensed_librispeech_asr.real_data_sd_asrkorean-asr
korean-asr — Korean ASR pseudo-labels for YODAS2
This repository contains transcripts and segment metadata only. It does not contain audio.
Every row points into espnet/yodas2 by
(shard, video_id, start, end), so you fetch the audio from the upstream dataset and cut it
yourself. See Reconstructing the audio.
split
utterances
hours
train
1,034,181
6,974.3
heldout
46,542
314.6
dev (subset of heldout)
3,000
20.4
Total labelled: 1,080,723 utterances / 7,288.9… See the full description on the dataset page: https://huggingface.co/datasets/nhatminh/korean-asr.persian-asr-text-2.69M-deduped
🗂️ persian-asr-text-2.69M-deduped
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Deduplicated Persian ASR text dataset used by the training stack.
پیکرهٔ متنی فارسیِ حذفتکرارشده برای ساخت واژگان، مدلسازی زبانی و پشتیبانی از آموزش ASR.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
4 files; approximately 109.64 MB
4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.russian-asr-leaderboardtajik-asr-youtube
tajik-asr-youtube
Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk
shows, podcasts, audiobooks, and learning content — with machine transcripts and the
verification scores left in as columns instead of applied as a filter. Pick your own
quality threshold; the training corpus this project actually ships
(tajik-asr-corpus-v3)
is the gated subset.
Layout
Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.indic-asr-benchmark
Indic ASR Benchmark — Nine Languages
Speech, human references, and side-by-side transcripts from multiple speech-to-text systems
across nine Indian languages — the evaluation data behind Navana's public Bodhi ASR
benchmark. Every clip comes from openly available research datasets, and every system is scored
the same way, on the same audio.
Released by Navana Tech, the team behind Bodhi, our
Indian-language speech-to-text engine.
A companion Hindi-only benchmark, with the full… See the full description on the dataset page: https://huggingface.co/datasets/Navana-AI/indic-asr-benchmark.aerograph-asrs
AeroGraph ASRS Dataset
2,000 real NASA Aviation Safety Reporting System (ASRS) incident reports
with LLM-extracted entities and relations for knowledge graph construction.
Dataset Description
This dataset contains processed ASRS incident narratives along with
structured entity and relation extractions conforming to an aviation
safety ontology (10 entity types, 8 edge types).
Reports Split
2000 reports from the NASA ASRS database
Fields: id, text, aircraft_type… See the full description on the dataset page: https://huggingface.co/datasets/Aryan95614/aerograph-asrs.
