datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nasle-mana-clean-chunked-30s
Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks
Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk.
Configuration
Rows
Columns
Meaning
labeled (train/)
9,886
audio, label
Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
Split
Rows
Audio
Columns
labeled
4,981
41.41 hours
audio, label
to_transcribe
11,127
92.72 hours
audio
The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.ganjoor-recitations-chunked
🗂️ ganjoor-recitations-chunked
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Ganjoor recitation chunked ASR dataset.
قطعههای تلاوت و خوانش گنجور برای آموزش و ارزیابی گفتار ادبی، شعر و خوانش رسمی فارسی.
🧩 Role
Persian speech dataset
مجموعهدادهٔ گفتار فارسی
📦 Snapshot
64 files; approximately 118.09 GB
64 فایل؛ حدود 118.09 GB
🧱 Packaging
61 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/ganjoor-recitations-chunked.asr-farsi-youtube-chunked-30-seconds
How To Use
from datasets import load_dataset
train = load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='train+val')
test =load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='test')
+300 Hours ASR dataset generated from this kaggle dataset
coraal_chunked
CORAAL Atlanta Dataset - Cleaned for ASR
Dataset Description
This dataset contains cleaned and segmented audio from the Corpus of Regional African American Language (CORAAL) Atlanta subset (v. 2020.05), specifically prepared for Automatic Speech Recognition (ASR) evaluation.
Dataset Summary
Source: CORAAL (Corpus of Regional African American Language) - Atlanta subset
Version: CORAAL:ATL v. 2020.05
Language: African American English (AAE)
Task: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/PardisTaghavi/coraal_chunked.ViMD_chunked_10s
ViMD Chunked 10s — 16kHz
Preprocessed from ViMD (Nguyen et al., EMNLP 2024).
Preprocessing
Resample: 44.1kHz -> 16kHz mono
Chunking: each audio is split into consecutive NON-OVERLAPPING segments
of at most 10 seconds. ALL segments are kept, including the final
remainder (no minimum length filter). 1 original file -> ceil(len/10s) samples.
Splits: original ViMD train/valid/test kept (speaker-exclusive).
Segments of the same file always stay in the same split.… See the full description on the dataset page: https://huggingface.co/datasets/tannhoo06/ViMD_chunked_10s.tbp_chunked_speech_restorised
tbp_chunked_speech_restorised
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised.cv_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/cv_chunked
Aligned dataset: instinct-org/cv_chunked_nfa_aligned
Rows: 71097 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_nfa_aligned.cv_chunked
cv_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked.default_voices_chunked_speech_restorised
default_voices_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised.cv_chunked_speech_restorised
cv_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_speech_restorised.default_voices_chunked_encoded
default_voices_chunked_encoded
This is a gated Uzbek encoded derived speech artifact from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_encoded.omni_chunked_speech_restorised
omni_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_speech_restorised.audiobook_chunked_speech_restorised
audiobook_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised.yt2_chunked_speech_restorised
yt2_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised.yt3_chunked_speech_restorised
yt3_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised.pods_chunked
Pods Chunked VAD Transcribed
This public, manually gated dataset contains VAD-produced speech chunks paired with transcripts.
Important: this release is not Sidonized yet. It has only completed VAD/chunk preparation and transcription. It has not gone through the later restoration, alignment, profiling, or final filtering stages.
Contents
Rows: 17966
Audio duration: 299.42 hours
Language: Uzbek
Access: public repository with manual gated access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/pods_chunked.tbp_chunked
tbp_chunked
This is a gated Russian chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked.yt1_chunked_speech_restorised
yt1_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised.audio_youtube_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/audio_youtube_chunked
Aligned dataset: instinct-org/audio_youtube_chunked_nfa_aligned
Rows: 559484 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_nfa_aligned.default_voices_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/default_voices_chunked
Aligned dataset: instinct-org/default_voices_chunked_nfa_aligned
Rows: 134236 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_nfa_aligned.miscellaneous_yt_chunked_speech_restorised_nfa_aligned
miscellaneous_yt_chunked_speech_restorised_nfa_aligned
Public, manually gated NFA-aligned Uzbek speech dataset derived from instinct-org/miscellaneous_yt_chunked_speech_restorised.
Contents
Parquet shards: 130
Rows: 528,187
Approx hours: 863.88
Audio column: audio with embedded FLAC bytes
Transcript column: transcription
Alignment columns: nfa_token_alignments, nfa_word_alignments, nfa_segment_alignments, nfa_character_alignments
Access And Use… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_nfa_aligned.default_voices_chunked
default_voices_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked.audio_youtube_chunked_speech_restorised
audio_youtube_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_speech_restorised.zy_chunked_speech_restorised
zy_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_speech_restorised.yt4_chunked_speech_restorised
yt4_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised.audiobook_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/audiobook_chunked
Aligned dataset: instinct-org/audiobook_chunked_nfa_aligned
Rows: 1291838 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_nfa_aligned.zy_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/zy_chunked
Aligned dataset: instinct-org/zy_chunked_nfa_aligned
Rows: 534816 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_nfa_aligned.tbp_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/tbp_chunked
Aligned dataset: instinct-org/tbp_chunked_nfa_aligned
Rows: 548483 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_nfa_aligned.
