datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
us_election_2024_telegram_distilled
A billion Telegram messages about the 2024 US presidential election
This is a dataset of Telegram messages collected during the 2024 US presidential election. For more details, see https://dl.acm.org/doi/10.1145/3701716.3715297.
~1.03B messages, ~43K chats, ~0.8TB (distilled).
~350M English messages have toxicity- and hate-related scores from the Perspective API. For more details, see https://support.perspectiveapi.com/s/about-the-api-attributes-and-languages?language=en_US.
~350M… See the full description on the dataset page: https://huggingface.co/datasets/leonardoblas/us_election_2024_telegram_distilled.russian-news-telegram-dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic News and Media,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the channel. views… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian-news-telegram-dataset.telegram-public-channels-2026-W36
Telegram public channels: a 5,177-channel snapshot with topics and a recommendation graph
A single snapshot of 5,177 public Telegram channels, measured on 31 August 2026, together with the
recommendation graph Telegram itself exposes between them.
Public Telegram data is hard to get in tabular form. The libraries that read it need a phone number
and a user session, and the two Telegram datasets that rank on Kaggle today are both from 2021.
This is a current measurement… See the full description on the dataset page: https://huggingface.co/datasets/starnikovoleg/telegram-public-channels-2026-W36.economic-telegram-news-corpus-2025
Economic Telegram News Corpus 2025
A corpus of 31,292 Russian-language economic news posts collected from 7 major Telegram channels, spanning January 2024 to September 2025. The dataset supports research on economic narrative detection, topic classification, and information diffusion in social media.
Associated Paper
Going Viral: LLM-Based Modeling of Economic Narratives
Dataset Description
The raw collection contains 123,273 posts. The economic corpus was… See the full description on the dataset page: https://huggingface.co/datasets/bruhwalkk/economic-telegram-news-corpus-2025.telegram-channel-dataset
Telegram AI Image Dataset — Cleaned for VLM LoRA Training
A cleaned dataset of 1,050 AI-generated images with their generation prompts, collected from a Chinese Telegram channel focused on GPT-Image-2 prompt engineering.
Each image is paired with a structured generation prompt in Chinese/English. Prompts have been cleaned of channel boilerplate — no bot instructions, hashtags, source credits, model name prefixes, emoji title lines, or channel footer ads.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GCStream/telegram-channel-dataset.cybersecurity_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic Cybersecurity,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the channel. views -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/cybersecurity_news_telegram_dataset.russian_oil_gas_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.telegram_ml_filtrationxenophobia_migrants_telegram
Anti-Immigrant Narrative Detection in Telegram News
Dataset Summary
This dataset contains Telegram messages from major news-oriented Telegram channels, collected to examine presence of negative and anti-immigrant narratives in media reporting compared to general crime reporting.
The dataset is intended for studying the prevalence, dynamics, and framing of anti-immigrant messaging in news sources, rather than detecting hate speech or legal violations.… See the full description on the dataset page: https://huggingface.co/datasets/aletheos-ngo/xenophobia_migrants_telegram.telegram-filtered-messagestelegram-financial-sentiment-summarizationtweets_about_german_politicians_jan_feb_2025_reddit_and_telegram
Dataset Card for LLM-based Detection of Manipulative Political Narratives
Dataset Summary
This dataset comprises an unfiltered collection of 1,255,895 short social media posts.
The language distribution is approximately 80% German and 20% English.
Supported Tasks and Leaderboards
text-classification
sentiment-analysis
Dataset Structure
Data Instances
A typical instance represents a single social media post referencing a specific… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_about_german_politicians_jan_feb_2025_reddit_and_telegram.tweets_about_german_politicians_jan_feb_2025_reddit_telegram_classified_embed_reduced_clustered
Dataset Card for LLM-based Detection of Manipulative Political Narratives (Embedded & Clustered)
Dataset Summary
This dataset represents the advanced analytical stage of a computational framework designed to process political content in social media.
It contains a filtered subset of 114,876 social media posts that were previously classified as containing manipulative strategic narratives (FIMI) by a prompt-driven reasoning model.
To facilitate narrative discovery without… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_about_german_politicians_jan_feb_2025_reddit_telegram_classified_embed_reduced_clustered.ru_telegramtweets_about_german_politicians_jan_feb_2025_reddit_and_telegram_classified
Dataset Card for LLM-based Detection of Manipulative Political Narratives (Classified)
Dataset Summary
This dataset represents a critical stage in a broader framework for processing political content in social media ecosystems.
It comprises an unfiltered collection of 1,255,895 short social media posts collected from X (formerly Twitter), Reddit, and Telegram.
The language distribution is approximately 80% German and 20% English.
The data was collected between January… See the full description on the dataset page: https://huggingface.co/datasets/SinclairSchneider/tweets_about_german_politicians_jan_feb_2025_reddit_and_telegram_classified.Telegram_Economic_Posts-RUjain-song-data-telegramPersian-Telegram-Conversations-50kiranian-telegram-news-messageswalle_reddit_telegram_with_sentiment_and_nerwalle_reddit_telegram_with_sentiment_and_ner_classifiedtelegram-channels-iranwalle_reddit_telegram_with_sentiment
