datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data-voice-vietnamese-restaurant-quan-oc
Vietnamese Restaurant Order Speech
This dataset contains Vietnamese spoken restaurant orders paired with text transcripts. Each utterance typically includes a table number, item quantities, dishes, drinks, and add-ons.
Dataset Structure
Files are split into subdirectories by filename-derived speaker_code to satisfy Hugging Face repository file-count limits:
metadata.csv: one row per audio sample.
audio/{speaker_code}/*.wav: mono WAV audio files.… See the full description on the dataset page: https://huggingface.co/datasets/EmilyNguyen235/data-voice-vietnamese-restaurant-quan-oc.2021-punctuation-restorationThis dataset is designed to be used in training models
that restore punctuation marks from the output of
Automatic Speech Recognition system for Polish language.kazakh-restaurantbe-sidon-restored-sample-100
be-sidon-restored-sample-100
Набор прыкладаў беларускай мовы з Common Voice (validated), апрацаваны мадэллю аднаўлення (denoise).
Файлы
audio/original/ — арыгінальныя кліпы (як у Common Voice).
audio/restored/ — адноўленыя WAV (48 кГц).
metadata.csv — палеткі: audio, original_audio, sentence, speaker.
Дата стварэння: 2025-09-25.
tbp_chunked_speech_restorised
tbp_chunked_speech_restorised
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised.default_voices_chunked_speech_restorised
default_voices_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised.cv_chunked_speech_restorised
cv_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_speech_restorised.audiobook_chunked_speech_restorised
audiobook_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised.omni_chunked_speech_restorised
omni_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_speech_restorised.yt2_chunked_speech_restorised
yt2_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised.miscellaneous_yt_chunked_speech_restorised_nfa_aligned
miscellaneous_yt_chunked_speech_restorised_nfa_aligned
Public, manually gated NFA-aligned Uzbek speech dataset derived from instinct-org/miscellaneous_yt_chunked_speech_restorised.
Contents
Parquet shards: 130
Rows: 528,187
Approx hours: 863.88
Audio column: audio with embedded FLAC bytes
Transcript column: transcription
Alignment columns: nfa_token_alignments, nfa_word_alignments, nfa_segment_alignments, nfa_character_alignments
Access And Use… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_nfa_aligned.audio_youtube_chunked_speech_restorised
audio_youtube_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_speech_restorised.zy_chunked_speech_restorised
zy_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_speech_restorised.education_chunked_speech_restorised
education_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked_speech_restorised.yt1_chunked_speech_restorised
yt1_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised.yt3_chunked_speech_restorised
yt3_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised.yt4_chunked_speech_restorised
yt4_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised.espeech_podcasts_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/espeech_podcasts_chunked_speech_restorised
Aligned dataset: instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned
Rows: 2467471 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned.yt_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt_chunked_speech_restorised
Aligned dataset: instinct-org/yt_chunked_speech_restorised_nfa_aligned
Rows: 416380 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised_nfa_aligned.espeech_podcasts_chunked_speech_restorised
espeech_podcasts_chunked_speech_restorised
This is a gated Russian speech-restorised chunked speech dataset from
instinct-org. It contains byte-backed speech audio and transcripts for
speech-to-text training, evaluation, forced alignment, and data preparation
workflows.
STT Alignment Source
Use this repository as the current repaired source candidate for Russian podcast
STT alignment. The matching base repository preserves the original transcript and
segment… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised.yt1_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt1_chunked_speech_restorised
Aligned dataset: instinct-org/yt1_chunked_speech_restorised_nfa_aligned
Rows: 261565 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised_nfa_aligned.yt2_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt2_chunked_speech_restorised
Aligned dataset: instinct-org/yt2_chunked_speech_restorised_nfa_aligned
Rows: 809612 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_nfa_aligned.yt3_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt3_chunked_speech_restorised
Aligned dataset: instinct-org/yt3_chunked_speech_restorised_nfa_aligned
Rows: 506444 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_nfa_aligned.yt4_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt4_chunked_speech_restorised
Aligned dataset: instinct-org/yt4_chunked_speech_restorised_nfa_aligned
Rows: 965192 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_nfa_aligned.yt_chunked_speech_restorised
yt_chunked_speech_restorised_48k
This is a gated Russian speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: ru (Russian)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised.miscellaneous_yt_chunked_speech_restorised
miscellaneous_yt_chunked_speech_restorised_48k
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised.
