Politics
Datasets
All datasets matching “Politics”news-politics-datasethumanual-politics
Humanual-Politics
Medium users responding to blog posts on political topics, featuring diverse political stances from users spanning different cultural backgrounds. This dataset is part of the HumanLM benchmark for training user simulators that accurately reflect real user behavior.Source: RapidAPI Medium endpoint · Domain: Long-form Content & Politics · Date Range: 2022-04-01 to 2025-11-04
The dataset contains 47,905 comments from 5,300 users across 14,724 posts, with an… See the full description on the dataset page: https://huggingface.co/datasets/snap-stanford/humanual-politics.task704_mmmlu_answer_generation_high_school_government_and_politics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task704_mmmlu_answer_generation_high_school_government_and_politics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task704_mmmlu_answer_generation_high_school_government_and_politics.IndustryCorpus_politics[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_politics.local_politicssports-politics-wikimedia
Sports vs Politics Wikipedia Dataset
A curated binary text classification dataset for distinguishing sports vs politics content, built from Wikipedia article extracts.
Dataset Description
Documents are extracted from Wikipedia using keyword-based retrieval and filtered using pattern matching to ensure clean class separation.
Collection Method
Keyword-based retrieval: For each class (sports/politics), 40 seed keywords were used to search Wikipedia and retrieve… See the full description on the dataset page: https://huggingface.co/datasets/Veeraraju/sports-politics-wikimedia.
