datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
legco-speech
香港立法會會議語音數據集
本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。
數據集製作流程
先去香港特別行政區立法會 YouTube下載所有會議紀錄並轉為 16kHz 採樣率嘅 OPUS音頻
用 fsmn-vad 切分所有語音,並用 Qwen3-ASR-1.7B 轉寫成粵文 srt 字幕
轉寫後用正則表達式修正字幕中常見轉寫錯誤
將數據集分成 raw、 segmented 兩個子集傳到HF
子集 subset
raw
segment
總行數 Row number
14,036
9,557,109
總時長 Total duration
22,195.55 hr (79,903,980.00 s)
20471.21 hr (73,696,365.27 s)
平均時長 Average duration
1.58 hr (5692.79 s)
7.71… See the full description on the dataset page: https://huggingface.co/datasets/laubonghaudoi/legco-speech.Speech2Latex
Speech2Latex Dataset
The Speech2LaTeX dataset is the first fully open-source large-scale dataset for converting spoken mathematical expressions and sentences into LaTeX. It comprises over 66,000 human-annotated audio samples of mathematical equations and sentences in both English and Russian, drawn from diverse scientific domains. This work lays the groundwork for future advances in multimodal AI, with a particular focus on mathematical content recognition.
The dataset was presented… See the full description on the dataset page: https://huggingface.co/datasets/marsianin500/Speech2Latex.reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.SpeechInstructBench
SpeechInstructBench
Arxiv: https://arxiv.org/abs/2503.02769
This is the SpeechInstructBench dataset download page.
SpeechInstructBench is a multilingual (Chinese and English) benchmark designed to evaluate the instruction-following capabilities of speech models. Instruction-following refers to a model’s ability to accurately interpret and execute user-provided natural language directives while strictly adhering to all specified constraints and requirements. To comprehensively… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/SpeechInstructBench.Hindi-speech-instruct
Hindi LLaMA-Omni Instruct Dataset
A Hindi speech instruction-following dataset designed for training speech-language models such as LLaMA-Omni. Each example pairs a spoken Hindi user question (audio) with a text assistant response.
Dataset Summary
Property
Value
Language
Hindi (hi)
Total examples
~110,718
Train split
~105,000 examples (batches 001–210)
Validation split
~5,500 examples (batches 211–222)
Audio format
FLAC, 16,000 Hz mono… See the full description on the dataset page: https://huggingface.co/datasets/Pastaaaaa2003/Hindi-speech-instruct.nigerian-pidgin-speech
Nigerian Pidgin Audio + Text Dataset for Whisper Fine-tuning
Nigerian Pidgin speech dataset for Whisper fine-tuning
Dataset Summary
This dataset contains audio recordings and transcriptions in Nigerian Pidgin English, designed for fine-tuning speech recognition models, particularly OpenAI's Whisper.
Dataset Structure
Train Split: 65 samples
Test Split: 8 samples
Total Duration: 0.0 hours (estimated)
Average Duration: 2.4 seconds per sample
Sample Rate: 16kHz… See the full description on the dataset page: https://huggingface.co/datasets/Rexe/nigerian-pidgin-speech.alpaca_speech_instruct
