CoolFace
Datasetpublic

sanjeevafk/great-speeches-corpus

๐ŸŽ™๏ธ Great Speeches Corpus: Definitive Historical Archives Benchmark (1893โ€“2015) Speaker Breakdown & Catalogue Index Namespace Historical Figure Time Span Speeches Word Count Character Count Primary Domain nehru Pandit Jawaharlal Nehru 1929โ€“1964 5,000 1,315,811 8,636,056 Nation Building, Democracy, NAM mandela Nelson Mandela 1951โ€“2013 1,000 1,141,540 7,083,904 Anti-Apartheid, Reconciliation, Human Rights kalam Dr. A. P. J. Abdul Kalam 1989โ€“2015โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/great-speeches-corpus.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes47downloads
Dataset Card

๐ŸŽ™๏ธ Great Speeches Corpus: Definitive Historical Archives Benchmark (1893โ€“2015)

![Hugging Face Dataset](https://huggingface.co/datasets/sanjeevafk/great-speeches-corpus) ![License: Apache 2.0](https://opensource.org/licenses/Apache-2.0) ![Dataset Viewer](https://huggingface.co/datasets/sanjeevafk/great-speeches-corpus)

Dataset Description

  • โ€”Curator: Sanjeevafk Speech Archives Project
  • โ€”Scope: 9 Iconic World Figures across Politics, Science, Philosophy, Technology, and Civil Rights
  • โ€”Time Span: 1893 โ€“ 2015 (122 Years)
  • โ€”Total Canonical Speeches: 8,687
  • โ€”Total Corpus Volume: 5,457,227 words (34,662,197 characters)
  • โ€”Dataset Hierarchy: Available as a unified dataset (config="all") or partitioned by speaker namespaces (config="kalam", config="nehru", config="mandela", etc.).

Speaker Breakdown & Catalogue Index

NamespaceHistorical FigureTime SpanSpeechesWord CountCharacter CountPrimary Domain
nehruPandit Jawaharlal Nehru1929โ€“19645,0001,315,8118,636,056Nation Building, Democracy, NAM
mandelaNelson Mandela1951โ€“20131,0001,141,5407,083,904Anti-Apartheid, Reconciliation, Human Rights
kalamDr. A. P. J. Abdul Kalam1989โ€“20159121,822,79411,374,971Vision 2020, Science, Youth Inspiration
ambedkarDr. B. R. Ambedkar1916โ€“195653774,465558,227Constitution of India, Social Justice, Law
vivekanandaSwami Vivekananda1893โ€“1902500123,276807,740Four Yogas, Parliament of Religions, Vedanta
einsteinAlbert Einstein1909โ€“195526064,815425,995Relativity, Nobel Lecture, Nuclear Disarmament
boseNetaji Subhas Chandra Bose1921โ€“194521551,032353,909Azad Hind, INA, Freedom Struggle
feynmanRichard P. Feynman1955โ€“1988162845,5265,300,276Quantum Electrodynamics, Caltech Lectures
steve_jobsSteve Jobs1976โ€“201110117,968121,119Apple Keynotes, Technology, Design
`all`Complete Unified Benchmark1893โ€“20158,6875,457,22734,662,197Multidisciplinary Master Corpus

Usage Examples

1. Load Complete Multi-Speaker Corpus

python
from datasets import load_dataset

# Stream the entire 8,687-speech historical benchmark
dataset = load_dataset("sanjeevafk/great-speeches-corpus", "all", split="train")
print(f"Total speeches loaded: {len(dataset)}")

2. Load Specific Historical Speaker Namespaces

python
from datasets import load_dataset

# Load Jawaharlal Nehru (5,000 speeches)
nehru_corpus = load_dataset("sanjeevafk/great-speeches-corpus", "nehru", split="train")

# Load Swami Vivekananda (500 discourses)
vivek_corpus = load_dataset("sanjeevafk/great-speeches-corpus", "vivekananda", split="train")

# Load Albert Einstein (260 scientific lectures)
einstein_corpus = load_dataset("sanjeevafk/great-speeches-corpus", "einstein", split="train")

Schema

  • โ€”speech_id (string): Canonical identifier (SPEECH-ENTITY-YYYY-NNN)
  • โ€”canonical_event_id (string): Unique event identifier
  • โ€”title (string): Title of speech, address, keynote, or lecture
  • โ€”date (string): Event date (DD.MM.YYYY)
  • โ€”year (int): Year of presentation
  • โ€”location (string): City, Country/State
  • โ€”venue (string): Hall, parliament, university, stadium, or convention center
  • โ€”period (string): Era classification tag
  • โ€”event_type (string): presidential_address, keynote, lecture, parliamentary_speech, historic_oration
  • โ€”speaker (string): Name of the historical speaker
  • โ€”tldr (string): Multi-sentence summary of core themes and historic significance
  • โ€”transcript (string): Full verbatim text in markdown format
  • โ€”source_url (string): Authoritative archive URL
  • โ€”rights_status (string): Transcript rights status
  • โ€”sha256 (string): Bit-level SHA-256 integrity hash
  • โ€”tags (list[string]): Semantic domain and topical tags