datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telegram-news-ua-dataset
Aisberg Telegram News UA
A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.us_election_2024_telegram_distilled
A billion Telegram messages about the 2024 US presidential election
This is a dataset of Telegram messages collected during the 2024 US presidential election. For more details, see https://dl.acm.org/doi/10.1145/3701716.3715297.
~1.03B messages, ~43K chats, ~0.8TB (distilled).
~350M English messages have toxicity- and hate-related scores from the Perspective API. For more details, see https://support.perspectiveapi.com/s/about-the-api-attributes-and-languages?language=en_US.
~350M… See the full description on the dataset page: https://huggingface.co/datasets/leonardoblas/us_election_2024_telegram_distilled.telegram-audiobook-chizzled
Telegram Persian Audiobook Chizzled
1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release
This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.Telegram-DBTelegram-Cleaned-DBTelegram-Cleaned-DBtelegram-financial-signalsv2
Financial Trading Signals Sentiment Dataset
Overview
This dataset contains 4,664 trading signals extracted from Telegram group chats, focused on financial instruments such as Forex pairs, commodities (Gold, Silver), stocks, and indices. Each signal is labeled with a sentiment value for use in financial sentiment analysis and machine learning applications.
Data Description
Each record represents a trading signal and includes fields for symbol, sentiment (both… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/telegram-financial-signalsv2.Telegram-Cleaned-DBTelegram-DatabaseTelegram-DBTelegram-Cleaned-DBTelegramtelegramTelegramGuard
🛡️ antispam.bot
The AI guardian that keeps your Telegram groups clean — and actually answers your questions.
Open source (MIT), zero config, self-hostable — the TelegramGuard project.
🌐 English · 中文 · Русский · Español · العربية · فارسی
Add it to your group (zero setup)
Add the official bot — no install, no config, no cost.
→ @iLangGuardBot
Search @iLangGuardBot on Telegram
Add it to your group
Give it admin — delete messages + ban users… See the full description on the dataset page: https://huggingface.co/datasets/i-Lang/TelegramGuard.Telegramtelegram-spam-hamTelegram-DatasetTelegramTelegramTelegram-DB-1telegramrussian-news-telegram-dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic News and Media,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the channel. views… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian-news-telegram-dataset.russian-telegram-chat-logs
russian-telegram-chat-logs
This dataset contains messages extracted from Telegram chat history, processed and ranked by their "semantic load" (information density).
Dataset Overview
The data is stored in Parquet format, which provides efficient storage and maintains data types. Each row represents a single message that has passed through several stages of cleaning and analysis.
Column
Type
Description
message
string
The cleaned text of the Telegram message… See the full description on the dataset page: https://huggingface.co/datasets/KvaytG/russian-telegram-chat-logs.telegram-public-channels-2026-W36
Telegram public channels: a 5,177-channel snapshot with topics and a recommendation graph
A single snapshot of 5,177 public Telegram channels, measured on 31 August 2026, together with the
recommendation graph Telegram itself exposes between them.
Public Telegram data is hard to get in tabular form. The libraries that read it need a phone number
and a user session, and the two Telegram datasets that rank on Kaggle today are both from 2021.
This is a current measurement… See the full description on the dataset page: https://huggingface.co/datasets/starnikovoleg/telegram-public-channels-2026-W36.telegram-spameconomic-telegram-news-corpus-2025
Economic Telegram News Corpus 2025
A corpus of 31,292 Russian-language economic news posts collected from 7 major Telegram channels, spanning January 2024 to September 2025. The dataset supports research on economic narrative detection, topic classification, and information diffusion in social media.
Associated Paper
Going Viral: LLM-Based Modeling of Economic Narratives
Dataset Description
The raw collection contains 123,273 posts. The economic corpus was… See the full description on the dataset page: https://huggingface.co/datasets/bruhwalkk/economic-telegram-news-corpus-2025.telegramtelegrampersian-telegram-corpus
PersianPoetryQuotes Dataset
Dataset Description
Dataset Summary
this is a collection of Persian messages extracted from various Telegram channels.
telegram
