CoolFace
Datasetpublic

laylaylo/global-leaders-discourses

WorldwideSpeechText Paper (placeholder) · GitHub (placeholder) · Dashboard Dataset Description WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 33 countries, covering 67 leaders, 47,750 unique speeches, and 71,677 total records including translations (1959–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions. Each… See the full description on the dataset page: https://huggingface.co/datasets/laylaylo/global-leaders-discourses.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
1likes51downloads
Dataset Card

WorldwideSpeechText

[Paper (placeholder)](placeholder) · [GitHub (placeholder)](placeholder) · [Dashboard](https://huggingface.co/spaces/laylaylo/global-leaders-discourses-dashboard)


Dataset Description

WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 33 countries, covering 67 leaders, 47,750 unique speeches, and 71,677 total records including translations (1959–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions.

Each record is a single speech text sourced from official governmental archives and public records. Where a source publishes a speech in multiple languages, each language version is stored as a separate row sharing the same base speech_id (distinguished by a language suffix); the total record count therefore exceeds the unique speech count. Non-speech content — news articles, parliamentary question-and-answer sessions, debate transcripts, and interviews — was removed through source-specific heuristic filtering, retaining only prepared or delivered address texts.

More information on dataset construction can be found in the accompanying paper.


Countries & Leaders

Countries were selected to represent all four regime types in the V-Dem Institute's Regimes of the World (RoW) framework — Closed Autocracy, Electoral Autocracy, Electoral Democracy, and Liberal Democracy — and to provide at least one example per World Bank region where a qualifying country could be sourced. Within each country, priority was given to long-tenured leaders: extended time in office yields richer longitudinal material and tends to correlate with well-maintained official speech archives. Coverage was opportunistically extended to other leaders accessible through the same source when the marginal collection effort was small.

Per-country regime classifications, speech distributions, and interactive visualisations are available in the dataset dashboard.

CountryLeaders (unique speeches)Languages
ArgentinaAlberto Fernández (662)<br>Cristina Fernández de Kirchner (558)<br>Javier Milei (192)<br>Mauricio Macri (663)<br>Néstor Kirchner (41)Spanish
BangladeshSheikh Hasina (35)Bengali, English
BelarusAleksandr Lukashenko (289)Belarusian, English, Russian
BrazilDilma Rousseff (851)<br>Jair Bolsonaro (391)<br>Luiz Inácio Lula da Silva (620)English, Portuguese, Spanish
CanadaMark Carney (70)English, French
ChileGabriel Boric Font (857)<br>José Antonio Kast (59)<br>Sebastián Piñera Echenique (897)Spanish
ChinaXi Jinping (890)Chinese, English
CubaFidel Castro (1,145)<br>Miguel Díaz-Canel Bermúdez (358)<br>Raúl Castro (84)Arabic, English, French, German, Italian, Portuguese, Russian, Spanish
DenmarkAnders Fogh Rasmussen (224)<br>Helle Thorning-Schmidt (77)<br>Lars Løkke Rasmussen (155)<br>Mette Frederiksen (123)<br>Poul Nyrup Rasmussen (116)Danish, English
EgyptAbdel Fattah el-Sisi (1,664)—
EritreaIsaias Afwerki (64)English
FranceEmmanuel Macron (1,309)<br>François Hollande (1,611)<br>Nicolas Sarkozy (1,199)English, French
GermanyAngela Merkel (1,506)<br>Frank Walter Steinmeier (963)<br>Friedrich Merz (156)<br>Horst Köhler (394)<br>Joachim Gauck (450)<br>Olaf Scholz (651)English, French, German, Russian, Ukrainian
HungaryViktor Orbán (887)English, French, German, Hungarian
IndiaNarendra Modi (3,030)Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Meitei, Odia, Tamil, Urdu
IranAli Hosseini Khamenei (3,452)English, Persian
IsraelBenjamin Netanyahu (1,252)<br>Naftali Bennett (240)<br>Yair Lapid (114)Arabic, English, Hebrew
JapanShinzo Abe (397)Japanese
JordanAbdullah II ibn Al-Hussein (354)Arabic, English
MalaysiaDato' Seri Anwar Ibrahim (519)<br>Dato' Sri Mohd Najib Abdul Razak (351)<br>Tan Sri Dato' Haji Muhyiddin Mohd Yassin (10)<br>Tun Abdullah Ahmad Badawi (407)<br>Tun Dr. Mahathir Mohamad (1,756)English, Malay
NetherlandsMark Rutte (213)Dutch, English, German
New ZealandBill English (11)<br>Jacinda Ardern (121)<br>John Key (74)English
North KoreaKim Jong Un (34)Chinese, English, Japanese, Korean, Russian, Spanish
PolandAndrzej Duda (2,376)Polish
RomaniaKlaus Iohannis (46)Romanian
RussiaVladimir Putin (1,194)English, Russian
RwandaPaul Kagame (742)English, French, Indonesian, Swahili
SenegalMacky Sall (29)English, French
SingaporeLee Hsien Loong (834)Chinese, English, Malay, Tamil
South AfricaCyril Ramaphosa (932)<br>Jacob Zuma (600)English
SpainJosé Luis Rodríguez Zapatero (1,614)<br>Mariano Rajoy (1,634)<br>Pedro Sánchez (1,595)English, Spanish
Sri LankaChandrika Bandaranaike Kumaratunga (84)English
TurkeyRecep Tayyip Erdoğan (2,744)English, French, Turkish
VenezuelaHugo Chávez (742)<br>Nicolás Maduro (38)Spanish

Languages

Speeches appear in the original language of delivery. Several sources publish speeches in multiple languages (original and official translations); all available language versions are included. The dataset contains texts in 35 languages.

Language co-occurrence across the dataset is visualised in the paper's language chord diagram (paper/visuals/language_chord.png).


Dataset Structure

Loading

python
from datasets import load_dataset

# All countries combined
ds = load_dataset("laylaylo/global-leaders-discourses", "all", split="train")

# Single country
ds = load_dataset("laylaylo/global-leaders-discourses", "france", split="train")

Each country is a separate subset (config); the all config is their union. Every config has a single train split.

Fields

FieldDescription
speech_idUnique speech identifier
speaker_nameName of the political leader
speech_dateDate of delivery (ISO 8601, where available)
language_isoISO 639 language code of the speech text
speech_titleTitle or headline of the speech
speech_textFull cleaned text of the speech
source_urlOriginal source URL
country_isoISO 3166-1 alpha-3 country code
speech_locationLocation where the speech was delivered
speech_topicsTopics or themes (where available)
additional_metadataAdditional structured metadata

Dataset Creation

Source Data

Speeches were collected from official governmental websites, presidential office archives, and recognised public records. Collection used country-specific HTML parsers adapted to each source's structure. No automated crawling across the open web was performed; all sources are official or widely recognised archival repositories.

Processing

After extraction, non-speech records were removed using source- and country-specific heuristics — primarily filtering on document type labels, URL patterns, and structural features of the HTML. No translation, alignment, or further annotation was applied. Texts are preserved as published by the source.


Intended Use

This dataset is intended for:

  • —Computational political science — studying how rhetoric varies across regime types, regions, and time
  • —Multilingual NLP — large-vocabulary, long-form political text across diverse scripts and languages
  • —Longitudinal analysis — tracking how individual leaders' language evolves over extended tenures

It is not intended as a complete archival record of any leader's speeches, nor as a source for real-time or journalistic use.


Considerations

Data Completeness

This dataset aggregates publicly available speeches from official archives. Coverage is necessarily incomplete: digital archives vary in quality, retrospective digitisation is uneven, and some speeches may not have been published online. We collected as comprehensively as each source permitted, but make no claim to exhaustiveness for any leader or date range, and do not represent this dataset as a new archive — it is a unified, computationally accessible aggregation of existing public records.

Copyright and Licensing Information

Speeches were collected from public governmental and official sources. The licensing terms for each source are listed below. Users are responsible for reviewing and complying with the terms of use of each source before using this dataset.

CountrySourceLegal / Copyright
Argentinacasarosada.gob.arTerms and Conditions
Argentinacfkargentina.com—
Bangladeshalbd.orgTerms and Conditions
Belaruspresident.gov.byAbout this page
Brazilbiblioteca.presidencia.gov.br—
Brazilgov.brTerms of Use
Canadapm.gc.caImportant Notices
Chileprensa.presidencia.cl—
Chinaen.cppcc.gov.cn—
Chinaen.people.cnLegal / Copyright Page
Chinajhsjk.people.cn—
Cubacuba.cu—
Cubacubadebate.cu—
Cubacubaminrex.cu—
Cubagranma.cu—
Cubajuventudrebelde.cu—
Cubapresidencia.gob.cu—
Denmarkenglish.stm.dk—
Denmarkstm.dk—
Egyptpresidency.eg—
Egyptsis.gov.eg—
Eritreashabait.com—
Franceelysee.frLegal Notice
Germanybundeskanzler.de—
Germanybundesregierung.deLegal Notice
Hungary2010-2014.kormany.huAbout this page
Hungary2015-2022.miniszterelnok.hu—
Hungaryminiszterelnok.hu—
Indiapmindia.gov.inWebsite Policies
Iranenglish.khamenei.irCreative Commons License)
Iranfarsi.khamenei.ir—
Irankhamenei.ir—
Israelgov.ilTerms of Use
Japanabeshinzo-digitalmuseum.com—
Japanen.abeshinzo-digitalmuseum.com—
Jordankingabdullah.joLegal Notice
Malaysiapmo.gov.myCopyright
Netherlandsgovernment.nlCopyright
Netherlandsopen.overheid.nl—
Netherlandsrijksoverheid.nl—
New Zealandbeehive.govt.nzCopyright
North Koreakcna.kp—
Polandprezydent.plCopyright
Romaniaarhiva.presidency.ro—
Romaniaklausiohannis.presidency.ro—
Romaniapresidency.roTerms and Conditions
Russiaen.kremlin.ruCopyright
Russiakremlin.ruCopyright
Rwandapaulkagame.rw—
Senegalmackysallfoundation-p2d.org—
Singaporepmo.gov.sgTerms of Use
South Africapresidency.gov.zaLegal Disclaimers
South Africathepresidency.gov.zaLegal Disclaimers
Spainlamoncloa.gob.esLegal Notice
Sri Lankapresidentcbk.lk—
Turkeytccb.gov.tr—
Venezuelamppre.gob.ve—
Venezuelatodochavez.gob.veAbout this page
Disclaimer: We do not guarantee the accuracy of the legal information above and take no responsibility for any use that conflicts with applicable licences or laws. Users are solely responsible for ensuring their use complies with the relevant terms of service and applicable law in their jurisdiction.

Limitations

  • —Speech counts vary substantially across leaders — from tens to tens of thousands — reflecting differences in source archive quality, a leader's total output, and tenure length rather than editorial selection within a collected archive.
  • —Temporal coverage is densest from approximately 2015 onward, where governmental online archives are most comprehensive.
  • —Some sources publish only selected speeches; completeness for a given leader depends entirely on what the source has made publicly available.

Citation

If you use this dataset, please cite:

bibtex
@misc{WorldwideSpeechText,
  author    = {Yaayladere, Leyla and Jin, Zhijing},
  title     = {WorldwideSpeechText: A Multilingual Corpus of Political Leader Speeches Across Regime Types},
  year      = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/laylaylo/global-leaders-discourses}}
}

Acknowledgements

We would like to thank Aman Gokrani, Berelian Karimian, and Miu Takagi for their contributions to data collection.