z-dickson/xlm-roberta-large-political-issues
Multilingual Political Issue Classifier
This model classifies political party press releases (or other similar political documents) according to the primary issue they address.
The classification scheme is very similar to the Comparative Agendas Project (see: https://www.comparativeagendas.net/), with the exception that the 'environmental' category includes climate change policies.
The model was trained on a dataset of over 140,000 examples drawn from four complementary sources: zero-shot labels generated by OpenAI's GPT-4.1-nano (30k press releases) and GPT-5-mini models (15k press releases); publicly available human-coded press release data from the Comparative Agendas Project and press releases coded by graduate research assistants (approximately 4k).
The issue scheme is as follows:
CATEGORY SUMMARIES
- Macroeconomics Covers broad economic policy topics like interest rates, inflation, unemployment, taxes, budgets, monetary and industrial policy. Also includes wage/price control and other macroeconomic matters.
- Civil Rights Focuses on discrimination (racial, gender, age, disability), voting rights, freedom of speech, privacy, and minority protections. Also includes anti-government groups and other civil rights topics.
- Health Encompasses healthcare reform, insurance, medical facilities and liability, workforce, and public health efforts. Covers topics from mental health and child health to drug abuse, R&D, and disease prevention.
- Agriculture Addresses farm subsidies, food safety, marketing, animal/crop disease, fisheries, and agricultural R&D. Also includes general agriculture policy and rural development.
- Labor Covers job safety, training, benefits, labor standards, unions, and youth/migrant employment. Also includes pensions and employment policies.
- Education Includes all education levels from early childhood to higher education, as well as special education, vocational training, and education quality initiatives. Also includes R&D and underserved student support.
- Environment and climate change Deals with water and air pollution, waste disposal, hazardous materials, conservation, endangered species, and indoor/outdoor environmental safety. Includes recycling, R&D, and land preservation.
- Energy Focuses on energy sources like nuclear, coal, oil, renewables, and electricity. Includes energy efficiency, conservation, and related R&D.
- Immigration Covers immigration laws, refugee policy, and citizenship issues.
- Transportation Addresses infrastructure, public transit, highways, air and rail travel, maritime transport, and transportation R&D.
- Law and Crime Includes crime control, enforcement, courts, prisons, drug crime, family law, juvenile justice, and terrorism. Also covers agencies, white-collar crime, and child abuse.
- Social Welfare Focuses on programs for low-income families, elderly and disabled assistance, child care, and volunteer organizations. Encompasses general welfare policies.
- Housing Covers public housing, community and rural development, housing for veterans, elderly, and the homeless. Includes affordability and urban planning.
- Domestic Commerce Includes banking, finance, small business, consumer protection, corporate governance, and commerce-related R&D. Also covers insurance, tourism, and bankruptcy.
- Defense Encompasses military policy, readiness, procurement, personnel, nuclear arms, foreign operations, and civil defense. Covers contractors, intelligence, and environmental compliance.
- Technology Covers space exploration, telecommunications, computing, broadcasting, and cybersecurity. Also includes scientific research, tech development, and commercial use of space.
- Foreign Trade Deals with trade agreements, tariffs, exports/imports, competitiveness, and exchange rates. Also includes international business and investment policy.
- International Affairs Includes diplomacy, foreign aid, developing countries, human rights, global organizations, and international finance. Also covers terrorism, embassies, and treaties.
- Government Operations Addresses bureaucracy, procurement, civil service, campaigns, tax enforcement, and census data. Also includes scandals, national holidays, and intergovernmental relations.
- Public Lands Covers parks, indigenous issues, forest and land management, water resources, and U.S. territories. Focuses on conservation and federal land use.
- Culture Encompasses general cultural policies, likely including funding, preservation, and promotion of cultural initiatives.
Countries/languages included in fine-tuning
The countries included are: Poland, Germany, Ireland, Netherlands, Slovenia, Denmark, Hungary, Austria, Sweden, Bulgaria, Spain, Croatia, Finland, United Kingdom, Greece, Switzerland, Estonia, France, Portugal, Cyprus, Slovakia, Italy, Czech Republic, and Belgium.
Accuracy
The model achieves a F1 score of 0.7958.
from transformers import AutoModelForSequenceClassification
from transformers import TextClassificationPipeline, AutoTokenizer
mp = 'z-dickson/xlm-roberta-large-political-issues'
model = AutoModelForSequenceClassification.from_pretrained(mp)
tokenizer = AutoTokenizer.from_pretrained(mp)
classifier = TextClassificationPipeline(tokenizer=tokenizer, model=model, device=0)
classifier("""
To ask the Secretary of State for Energy and Climate \\
Change what estimate he has made of the proportion of carbon \\
dioxide emissions arising in the UK attributable to burning.
"""
)IDX_TO_CAP = {
0: 1, # Macroeconomics
1: 2, # Civil Rights
2: 3, # Health
3: 4, # Agriculture
4: 5, # Labor
5: 6, # Education
6: 7, # Environment & Climate Change
7: 8, # Energy
8: 9, # Immigration
9: 10, # Transportation
10: 12, # Law and Crime
11: 13, # Social Welfare
12: 14, # Housing
13: 15, # Domestic Commerce
14: 16, # Defense
15: 17, # Technology
16: 18, # Foreign Trade
17: 19, # International Affairs
18: 20, # Government Operations
19: 21, # Public Lands
20: 23, # Culture
}Citation
The data collection efforts of the press releases were originally from the following work. The two lead authors created the model in additional collaboration.
@article{KRIIT2026,
title = {The 2026 Europe report of the Lancet Countdown on health and climate change: narrowing window for decisive health action},
journal = {The Lancet Public Health},
year = {2026},
issn = {2468-2667},
doi = {https://doi.org/10.1016/S2468-2667(26)00025-3},
url = {https://www.sciencedirect.com/science/article/pii/S2468266726000253},
author = {Hedi K Kriit and José Chen-Xu and Jan C Semenza and Hannah Heiliger and Anil Markandya and Niheer Dasandi and Slava Jankin and Kim R {van Daalen} and Hicham Achebak and Anna Alari and Tilly Alcayna and Emily Ball and Joan Ballester and Hannah Bechara and Max W Callaghan and Monique {van Cauwenberghe} and Gina E C Charnley and Orin Courtenay and Marta Cirach and Paulina Garcia-Corral and Troy J Cross and Shouro Dasgupta and Zachary P Dickson and Matthew J Eckelman and Cornelius Erfort and Peter Fransson and Zia Farooq and Olga Gasparyan and Ian Hamilton and Marlies Hesselman and Risto Hänninen and Shih-Che Hsu and Tomáš Janoš and Harshavardhan Jatkar and Ollie Jay and Harry Kennard and Kajal Khanna and Gregor Kiesewetter and Rachel Lowe and Daniela Lührsen and Carla Maia and Jaime Martinez-Urtaza and Jan C Minx and Mark Nieuwenhuijsen and Julia Palamarchuk and Adria {San José Plana} and Tim Repke and Jorge A Roa-Contreras and Elizabeth J Z Robinson and Daniel Scamman and Natalia Shartova and Jodi D Sherman and Elena Sirotkina and Pratik Singh and Mikhail Sofiev and Marco Springmann and Lara Stucki and Federico Tartarini and Joaquin Triñanes and Maria Walawender and Marina Romanello and Josep M Antó and Maria Nilsson and Cathryn Tonne and Joacim Rocklöv}
}Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 2e-05
- trainbatchsize: 64
- evalbatchsize: 64
- seed: 42
- gradientaccumulationsteps: 2
- totaltrainbatch_size: 128
- optimizer: Use OptimizerNames.ADAMWTORCHFUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
- lrschedulertype: cosine
- lrschedulerwarmup_steps: 0.1
- num_epochs: 10
