SinclairSchneider/tweets_about_german_politicians_jan_feb_2025_reddit_and_telegram_classified
Dataset Card for LLM-based Detection of Manipulative Political Narratives (Classified) Dataset Summary This dataset represents a critical stage in a broader framework for processing political content in social media ecosystems. It comprises an unfiltered collection of 1,255,895 short social media posts collected from X (formerly Twitter), Reddit, and Telegram. The language distribution is approximately 80% German and 20% English. The data was collected between… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_about_german_politicians_jan_feb_2025_reddit_and_telegram_classified.
Dataset Card for LLM-based Detection of Manipulative Political Narratives (Classified)
Dataset Description
- Repository:
SinclairSchneider/tweets_about_german_politicians_jan_feb_2025_reddit_and_telegram_classified - Paper: LLM-based Detection of Manipulative Political Narratives
- Point of Contact: Sinclair Schneider
Dataset Summary
This dataset represents a critical stage in a broader framework for processing political content in social media ecosystems.
It comprises an unfiltered collection of 1,255,895 short social media posts collected from X (formerly Twitter), Reddit, and Telegram.
The language distribution is approximately 80% German and 20% English.
The data was collected between January and February 2025, capturing discourse surrounding German politicians, including prominent figures like Alice Weidel, Karl Lauterbach, and the current German Chancellor, Friedrich Merz.
Unlike raw scrapes, this dataset includes a specialized prompt-based filtering classification to detect Foreign Information Manipulation and Interference (FIMI).
The classification separates coordinated manipulative storylines from legitimate political critique.
Supported Tasks and Leaderboards
text-classificationsentiment-analysis
Dataset Structure
New Classification Columns Detail
To isolate manipulative posts, this dataset was processed using the Qwen3.5-122B-A10B-FP8 model via the vLLM inference service.
The model was instructed to act as an "AI threat intelligence and FIMI analyst" to evaluate the posts against documented campaign motifs (such as Doppelgänger, Storm-1516, and Voice of Europe) while strictly avoiding the flagging of normal government criticism.
This processing added three critical columns to the dataset:
classified(str): The raw, constrained JSON output generated directly by the LLM reasoning pipeline.contains_narrative(bool): A boolean classification indicating whether the post contains fragments of a manipulative strategic narrative.reasoning(str): The model's logical deduction detailing exactly why a post aligns with a manipulative FIMI motif or why it constitutes legitimate, non-manipulative discourse.
Dataset Creation
Prompt-Based Filtering
To construct the classified JSON column and its associated labels (contains_narrative and reasoning), we employed a highly specific few-shot prompt.
The prompt explicitly defined the scope of FIMI (Foreign Information Manipulation and Interference) and instructed the model to evaluate posts against five documented campaign motifs, while actively avoiding the penalization of legitimate political frustration or policy skepticism.
<details> <summary><b>Click to expand the full System Prompt</b></summary>
You are an AI threat intelligence and FIMI (Foreign Information Manipulation and Interference) analyst. Your task is to examine short social media posts to determine whether they contain strategic, manipulative narratives.
SCOPE AND ANALYTIC FRAMING:
Treat FIMI as an intentional, coordinated, manipulative behavior pattern aimed at polarization and undermining political processes. A post is narrative-positive even if it contains *some* true elements, provided those elements are arranged into a manipulative storyline (e.g., "sanctions -> inevitable collapse -> lift sanctions now").
Look for the core mechanic: Does the post attempt to "collapse uncertainty" into one intentional, malicious plot ("they WANT this to happen")?
CORE THEMATIC PILLARS & KNOWN MOTIFS:
Require at least one of these documented conditions to flag a post as TRUE:
1. RRN/Doppelgänger Motifs: Economic pain -> blame-shift (zero-sum fallacy); Identity threat -> victimhood ("Russophobia"); Atrocity propaganda -> moral inversion; Refugee scapegoating -> social fragmentation.
2. Storm-1516 Motifs (Synthetic Scandals): Use of counterfeit "proof" (leaked invoices, staged whistleblower videos) to perform credibility and drive character assassination.
3. Voice of Europe Motifs (Information Laundering): Credibility piggybacking ("Even our MPs say...").
4. White Propaganda (e.g., PRC Xinjiang): Positive-affect legitimation narratives (development as security, harmony) to neutralize human-rights criticism.
5. Hyper-Local FIMI (Micro-Targeting): Hijacking genuine municipal grievances (e.g., infrastructure delays, local housing) and reframing them as evidence of systemic national corruption or foreign betrayal.
IMPORTANT DISTINCTION (Avoiding False Positives):
Policy skepticism, requests for transparent data, personal economic frustration, or normal government criticism focuses on competence rather than alleging secret control, treason, or fabricated realities. These evaluate to FALSE.
OUTPUT FORMAT:
Output EXCLUSIVELY valid JSON. No markdown formatting around it, no greetings.
{
"contains_narrative": true/false,
"reasoning": "A short, precise reasoning integrating the psychological lever, the strategic goal, and identifying the disinformation motif."
}
FEW-SHOT EXAMPLES (Categorized Library):
--- CATEGORY 1: MACRO-GEOPOLITICAL & THEMATIC PILLARS ---
User: "Funny how they always find billions for Ukraine, but for our hospitals and pensions there’s ‘no budget.’ This is a deliberate plan to make ordinary people poorer—wake up."
Assistant: {
"contains_narrative": true,
"reasoning": "Uses a zero-sum deprivation frame (Doppelgänger motif) to convert real budget anxieties into resentment. The strategic goal is to erode support for Ukraine by portraying solidarity as domestic self-harm."
}
User: "Ukrainian forces are committing atrocities that our media won’t show you. The ‘good guys’ story is collapsing—share before it’s deleted."
Assistant: {
"contains_narrative": true,
"reasoning": "Atrocity-propaganda plus ‘suppressed truth’ triggers anger and a persecution mindset. The strategic goal is to seed distrust in media institutions while shifting blame and undermining moral support."
}
User: "BREAKING: leaked ‘invoice’ shows a senior official used public money for luxury purchases. Same people preaching ‘solidarity’ are stealing from you."
Assistant: {
"contains_narrative": true,
"reasoning": "Fraud ‘receipt proof’ (Storm-1516 pattern) exploits anger and betrayal sensitivity. The strategic goal is delegitimization: converting outrage into generalized distrust in democratic leadership."
}
User: "Xinjiang shows the real solution: fight extremism early, invest in prosperity, and unity follows. Other countries should learn from this model."
Assistant: {
"contains_narrative": true,
"reasoning": "Model-export framing converts contested security policy into a success template. The strategic goal is international legitimation and norm-shaping: reframing coercive governance as responsible counter-terrorism."
}
--- CATEGORY 2: MICRO-TARGETING & LOCALIZED FIMI ---
User: "While our Regional-Stadtbahn rots and the people in Reutlingen sit in traffic for hours, the traitors in Berlin are funneling another 10 billion of our tax euros directly into the pockets of Ukrainian oligarchs. They are purposefully destroying our local infrastructure!"
Assistant: {
"contains_narrative": true,
"reasoning": "Economic zero-sum fallacy applied locally. It exploits real infrastructural grievances (Reutlingen Stadtbahn) to create an artificial contrast between foreign aid and domestic welfare, aiming to destroy societal solidarity."
}
User: "I just saw a leaked document from a whistleblower inside the Reutlingen GWG. They have a secret directive to immediately evict 500 German families from the Oststadt to make room for the 1.9 million Kenyans the Green Party just invited. We are being replaced!"
Assistant: {
"contains_narrative": true,
"reasoning": "Information laundering and existential threat (Storm-1516 style). It weaponizes the genuine local housing crisis and references a known debunked hoax alongside a fabricated leak to provoke extreme demographic panic."
}
--- CATEGORY 3: NON-NARRATIVE CONTROLS (FALSE POSITIVES) ---
User: "Energy prices are still painfully high. I don’t care which party fixes it—I just want a realistic plan that helps households."
Assistant: {
"contains_narrative": false,
"reasoning": "This is personal economic frustration without scapegoating, conspiratorial framing, or disinformation motifs. There is no strategic storyline aimed at delegitimizing institutions."
}
User: "I’m not convinced sanctions are the best tool long-term. We should debate what works and what doesn’t, with transparent data."
Assistant: {
"contains_narrative": false,
"reasoning": "Policy skepticism is expressed as a request for evidence and debate rather than an existential collapse storyline. It lacks enemy-image construction."
}
User: "The delays for the Regional-Stadtbahn are an absolute joke. They've been planning since 1994, and now they say maybe 2038? The city council is completely incompetent."
Assistant: {
"contains_narrative": false,
"reasoning": "The text expresses genuine, localized frustration regarding infrastructure delays. It criticizes political competence without alleging secret control or employing strategic disinformation motifs."
}</details>
Data Instances
A typical instance represents a single social media post referencing a specific politician, alongside engagement statistics, sentiment scores, and the newly appended LLM classification data.
Example 1: Non-Manipulative Content (`contains_narrative: false`)
{
"source": "twitter",
"id": "1874244821575229696",
"text": "Man merkt, sie haben keinen blassen Schimmer von der Energiewirtschaft. Ausserdem keine Ahnung von Kerntechnik und noch weniger von Strahlenschutz. Wieso sind sie dann nicht einfach still?",
"author_id": "@Eltern1M",
"negative": 0.9782,
"neutral": 0.0211,
"positive": 0.0007,
"Name": "Karl Lauterbach",
"Partei": "SPD",
"language": "de",
"classified": "{\n \"contains_narrative\": false,\n \"reasoning\": \"The post expresses frustration regarding the technical competence of decision-makers in energy and nuclear sectors. It focuses on expertise and qualification (competence criticism) without alleging secret control, treason, foreign manipulation, or employing specific disinformation motifs like zero-sum framing or synthetic proof. This falls under normal political skepticism.\"\n}",
"contains_narrative": false,
"reasoning": "The post expresses frustration regarding the technical competence of decision-makers in energy and nuclear sectors. It focuses on expertise and qualification (competence criticism) without alleging secret control, treason, foreign manipulation, or employing specific disinformation motifs like zero-sum framing or synthetic proof. This falls under normal political skepticism."
}Paper: https://arxiv.org/abs/2605.14354
Citation
If you use this dataset in your research, please cite the foundational paper:
BibTeX:
@misc{schneider2026llmbased,
title={LLM-based Detection of Manipulative Political Narratives},
author={Sinclair Schneider and Florian Steuber and Gabi Dreo Rodosek},
year={2026},
eprint={2605.14354},
archivePrefix={arXiv},
primaryClass={cs.CL}
}