datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hsk-sentences-audio
HSK Sentences Audio
4,354 Chinese sentences graded against the official HSK 3.0 levels 1–6, with
pinyin, English translations, per-word glosses, grammar tags, and normal/slow
synthetic speech. The complete export contains 8,708 MP3 files.
Dataset structure
The Viewer reads native Parquet from data/train.parquet, avoiding a dependency
on Hugging Face's JSON-to-Parquet conversion service. The same 4,354 records are
also available as validated JSON Lines in… See the full description on the dataset page: https://huggingface.co/datasets/no7z/hsk-sentences-audio.gsat-vocab-sentences-tts
GSAT Vocabulary TTS Audio
Text-to-speech audio files for GSAT (General Scholastic Ability Test) English vocabulary.
Structure
audio/ - MP3 audio files organized by hash prefix (e.g., audio/ab/abcd1234....mp3)
index.jsonl - Index file mapping hashes to text and TTS engine used
Engines
Kokoro (af_heart voice) - Used for lemmas (single words/phrases)
Supertonic (M1 voice) - Used for example sentences
Audio Format
Format: MP3
Sample rate: 24kHz… See the full description on the dataset page: https://huggingface.co/datasets/TCabbage/gsat-vocab-sentences-tts.Taiwanese-Minnan-Example-Sentences
Taiwanese Minnan Example Sentences
The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems.
Dataset Features
Source: Ministry of Education, Taiwan (Sutian Resource Center)
Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.mixed-language-detection-pilot-complete-sentences
Mixed-Language Speech Detection Pilot — Complete Sentences
This is the complete-sentence revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed
Earlier generation forced source transcripts into arbitrary… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-complete-sentences.visualears-hardword-sentences
🗂️ visualears-hardword-sentences
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Hard-word sentence dataset used for semantic/keyword stress cases.
جملههای دارای واژههای دشوار و معنایی برای آزمون تنش واژگانی، بازیابی کلیدواژه و S³.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
266 files; approximately 39.23 GB
266 فایل؛ حدود 39.23… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-hardword-sentences.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.ears_dataset_sentencesswiss-german-city-sentences_trainswiss-german-city-sentences_valswiss-german-city-sentences_v2
Swiss German City Sentences v2
Synthetic Swiss German speech dataset with city name sentences across multiple dialects.
English-Hebrew-Mixed-Sentences
English-Hebrew Mixed Sentences Dataset
A dataset of English sentences with Hebrew words and phrases interspersed, designed for speech-to-text training and evaluation for English speakers in Israel.
Overview
This dataset addresses a common challenge for English-speaking immigrants in Israel: standard speech-to-text (STT) systems struggle to accurately transcribe code-switched speech where Hebrew words are mixed into primarily English sentences.
Example: "I need to pick up… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/English-Hebrew-Mixed-Sentences.ASR_Marathi_Sentences
Marathi Sentence-Level ASR Dataset
📌 Overview
This dataset contains sentence-level Marathi speech segments aligned with transcripts.
The dataset was created by extracting subtitle timestamps (SRV3 format) from Marathi YouTube content and segmenting the corresponding audio using precise time alignment.
Each sample contains:
A WAV audio file (sentence-level)
The corresponding Marathi transcript text
This dataset is suitable for:
Whisper fine-tuning
Wav2Vec2 CTC training… See the full description on the dataset page: https://huggingface.co/datasets/transitionGap/ASR_Marathi_Sentences.child_handpicked_sentencesASR_Hanyanvi_7k_sentencesASR_Haryanvi_1K_sentenceschild_handpicked_sentencesASR_1kBhojpuri_Sentencessynthetic_utterances_pt-BR_1400_sentences_x_6_speakerscamara_audio_sentencesASR_1kHindi_Sentencessynthetic_utterances_pt-BR_1200_sentences_x_6_speakers30_report_sentences_datasettts_synthetic_kn_single_sentencesASR_Marathi_Sentences
Marathi Sentence-Level ASR Dataset
📌 Overview
This dataset contains sentence-level Marathi speech segments aligned with transcripts.
The dataset was created by extracting subtitle timestamps (SRV3 format) from Marathi YouTube content and segmenting the corresponding audio using precise time alignment.
Each sample contains:
A WAV audio file (sentence-level)
The corresponding Marathi transcript text
This dataset is suitable for:
Whisper fine-tuning
Wav2Vec2 CTC training… See the full description on the dataset page: https://huggingface.co/datasets/Prasad12344321/ASR_Marathi_Sentences.eg-ADI-sentencesharvard-sentences-kokorokokoro-harvard-sentencesASR_Marathi_Sentences_1kS2T_SplitEndMovie_SentencesS2T_Merged_Sentences
