datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpeechInstructBench
SpeechInstructBench
Arxiv: https://arxiv.org/abs/2503.02769
This is the SpeechInstructBench dataset download page.
SpeechInstructBench is a multilingual (Chinese and English) benchmark designed to evaluate the instruction-following capabilities of speech models. Instruction-following refers to a model’s ability to accurately interpret and execute user-provided natural language directives while strictly adhering to all specified constraints and requirements. To comprehensively… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/SpeechInstructBench.speech-CMMLUThis dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
Usage
python audio_evals/main.py --dataset speech-cmmlu --model MiniCPMo2_6-speech --use_model_pool --workers 2
@article{ultraevalaudio,
title={UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models},
author={Qundong Shi and Jie Zhou and Biyuan Lin and Junbo Cui and Guoyang Zeng and Yixuan Zhou and… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/speech-CMMLU.steve-jobs-speech-corpus
🍏 Steve Jobs Lifetime Keynotes, Speeches & Interviews Corpus (1976–2011)
Historical Eras Distribution
Early Apple Era (1976–1985): 14 keynotes & speeches
NeXT & Pixar Wilderness Era (1985–1996): 16 product launches & oral histories
Apple Renaissance & Mac OS X Era (1997–2006): 51 landmark keynotes & interviews
The Mobile & Cloud Revolution (2007–2011): 20 revolutionary product introductions & final discourses
Key Landmark Ingests
1980 McKenna… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/steve-jobs-speech-corpus.german-parliament-speeches
German Parliament Speeches
This dataset contains speeches from the German parliament, derived from the Open Discourse Project (Harvard Dataverse).
Source
Data source:
Open Discourse ProjectHarvard DataverseDOI: 10.7910/DVN/FIKIBO
Original citation:
@data{DVN/FIKIBO_2020,
author = {Richter, Florian and Koch, Philipp and Franke, Oliver and Kraus, Jakob and Kuruc, Fabrizio and Thiem, Anja and Högerl, Judith and Heine, Stella and Schöps, Konstantin},
publisher = {Harvard… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/german-parliament-speeches.llama-omni-speech-instruct
Llama3.2 Omni Speech Instruct Dataset
This dataset is created for the sole purpose of enhancing the LLM capability to become multi-modals. This dataset has speech instruction
that a model could use to learn and produce the output thus allowing the model to overcome only text input and extends it capabilities
towards processing speech command as well.
Dataset Details
Dataset Description
This dataset can be used to train an LLM model to allow adaptibility in… See the full description on the dataset page: https://huggingface.co/datasets/gruhit-patel/llama-omni-speech-instruct.einstein-speech-corpus
🌌 Albert Einstein Lifetime Scientific Lectures & Pacifist Speeches Corpus (1909–1955)
Historical Eras Distribution
The Miracle Year & General Relativity (1909–1920): 60 foundational lectures & academy addresses
Nobel Prize & Global Relativity Lectures (1921–1932): 80 Princeton lectures, Nobel oration, Solvay debates, and world tours
Princeton IAS & Anti-Fascism (1933–1945): 60 Royal Albert Hall farewell, IAS seminars, and Roosevelt atomic letter
Nuclear… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/einstein-speech-corpus.dolly-15k-pirate-speechDataset for writing style transfer experimentation based on article:
https://ai-r.com/blog/pirate-linguistics-and-tone-of-voice-fine-tuning-llms-to-talk-like-swashbucklers
Only responses are in 'pirate speech'
arrr python library was used to simply change original responses to 'pirate speech' responses
https://pypi.org/project/arrr/
ambedkar-speech-corpus
⚖️ Dr. B. R. Ambedkar Lifetime Speeches, Constituent Assembly Debates & Writings Corpus (1916–1956)
A clean, standardized, machine-readable dataset comprising the landmark speeches, Constituent Assembly Debates (CAD), economic treatises, and philosophical discourses of Babasaheb Dr. Bhimrao Ramji Ambedkar (1891–1956), the chief architect of the Constitution of India.
📊 Dataset Statistics
Metric
Value
Total Canonical Speeches & Debates
537
Total… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/ambedkar-speech-corpus.mandela-speech-corpus
🇿🇦 Nelson Mandela Lifetime Speech & Lecture Corpus (1951–2010)
A comprehensive, clean, verified, machine-readable dataset comprising the complete lifetime speeches, parliamentary addresses, international keynotes, university convocations, and trial statements delivered by Nelson Rolihlahla Mandela across 6 decades (1951–2010).
📊 Dataset Statistics
Metric
Value
Total Canonical Speeches
1,000
Total Raw Source Records
1,000
Total Word Count
1… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/mandela-speech-corpus.nehru-speech-corpus
🏛️ Pandit Jawaharlal Nehru Lifetime Speeches, Parliamentary Debates & Treatises Corpus (1929–1964)
Historical Eras Distribution
Freedom Struggle & Congress Presidency (1929–1945): 1000 presidential addresses & anti-imperialist manifestos
Independence, Constituent Assembly & Partition (1946–1950): 1200 Tryst with Destiny, Objectives Resolution, Gandhi eulogy, and Red Fort orations
Nation Building, Five-Year Plans & Panchsheel (1951–1959): 1600 Temples of… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/nehru-speech-corpus.great-speeches-corpus
🎙️ Great Speeches Corpus: Definitive Historical Archives Benchmark (1893–2015)
Speaker Breakdown & Catalogue Index
Namespace
Historical Figure
Time Span
Speeches
Word Count
Character Count
Primary Domain
nehru
Pandit Jawaharlal Nehru
1929–1964
5,000
1,315,811
8,636,056
Nation Building, Democracy, NAM
mandela
Nelson Mandela
1951–2013
1,000
1,141,540
7,083,904
Anti-Apartheid, Reconciliation, Human Rights
kalam
Dr. A. P. J. Abdul Kalam
1989–2015… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/great-speeches-corpus.hausa-pq-speech-validated
Hausa WAEC PQ Speech Dataset (Validated)
Dataset Description
This dataset contains 1,361 multiple-choice WAEC past questions translated from English into Hausa, with human-validated Hausa translations. The data is designed to support speech synthesis, machine translation evaluation, and low-resource NLP research for Hausa — one of the most widely spoken languages in West Africa.
Languages
Source: English (en)
Target: Hausa (ha)… See the full description on the dataset page: https://huggingface.co/datasets/honourjesus/hausa-pq-speech-validated.feynman-speech-corpus
⚛️ Richard P. Feynman Complete Physics Lectures & Lifetime Keynotes Corpus (1955–1988)
A complete, machine-readable dataset comprising The Feynman Lectures on Physics (Volumes I, II, and III delivered at Caltech) along with his foundational landmark scientific keynotes, Cornell Messenger Lectures, Nobel address, and Rogers Commission reports.
📊 Dataset Statistics
Metric
Value
Total Canonical Lectures & Keynotes
162
Total Raw Source Records
162… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/feynman-speech-corpus.subhas-chandra-bose-speech-corpus
🇮🇳 Netaji Subhas Chandra Bose Lifetime Speeches & Revolutionary Addresses Corpus (1921–1945)
Historical Eras Distribution
Early Nationalist Movements & European Exile (1921–1937): 100 speeches & declarations
Congress Presidency & Forward Bloc (1938–1940): 45 presidential addresses & manifestos
Azad Hind & Singapore Proclamation (1941–1943): 9 sovereign proclamations & military orders
Military Campaign & Final Testaments (1944–1945): 61 battlefront addresses… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/subhas-chandra-bose-speech-corpus.vivekananda-speech-corpus
🕉️ Swami Vivekananda Complete Works & World Parliament of Religions Corpus (1893–1902)
Historical Eras Distribution
Parliament of Religions & Early American Tour (1893–1894): 125 discourses & orations
The Four Yogas & Western Mission (1895–1896): 125 yoga treatises & Harvard address
Lectures from Colombo to Almora & The Mission (1897–1898): 125 triumphant national awakenings & Ramakrishna Mission foundation
Second Western Visit & Final Testaments (1899–1902):… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/vivekananda-speech-corpus.kalam-speech-corpus
Dr. A.P.J. Abdul Kalam Lifetime Speech & Lecture Corpus
Dataset Summary
The Dr. A.P.J. Abdul Kalam Lifetime Speech & Lecture Corpus is a comprehensive, machine-readable dataset of speeches, addresses, convocations, and scientific lectures delivered by Dr. A.P.J. Abdul Kalam across his career as a rocket scientist, 11th President of India (2002–2007), and global statesman (2007–2015).
Total Canonical Events: 912 unique speech events
Total Raw Sources Tracked: 1… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/kalam-speech-corpus.spontanous-speech-qa
Spontanous speech QA
This dataset contains QA pairs from the spontaneous speech subsection of the Danish Gigaword.
The dataset is created from the DDSC dataset and
filtered to only include QA pairs where the question is less than 20 tokens and the answer is
at least 4 tokens long.
To find out more about the creation see the accompanying script.
speechllm-gaslighting-benchmark
Speech LLM Gaslighting Negation Benchmark
This dataset packages 5 speech tasks into a unified Hugging Face friendly format for studying gaslighting-style negation prompts against speech-language models.
It includes:
MELD: emotion classification
MMAU: audio reasoning
MMSU: spoken multiple-choice QA
OpenBookQA: spoken multiple-choice QA
VocalSound: vocal sound classification
All audio is normalized to 16 kHz. Each example contains:
one clean prompt
5 gaslighting prompt variants
the… See the full description on the dataset page: https://huggingface.co/datasets/Jack-ppkdczgx/speechllm-gaslighting-benchmark.
