datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telegram-audiobook-chizzled
Telegram Persian Audiobook Chizzled
1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release
This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.Telegram-DBTelegram-Cleaned-DBTelegram-Cleaned-DBTelegram-Cleaned-DBTelegram-DatabaseTelegram-DBtelegram-financial-signalsv2
Financial Trading Signals Sentiment Dataset
Overview
This dataset contains 4,664 trading signals extracted from Telegram group chats, focused on financial instruments such as Forex pairs, commodities (Gold, Silver), stocks, and indices. Each signal is labeled with a sentiment value for use in financial sentiment analysis and machine learning applications.
Data Description
Each record represents a trading signal and includes fields for symbol, sentiment (both… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/telegram-financial-signalsv2.Telegram-Cleaned-DBTelegramTelegramTelegramTelegram-DatasetTelegramTelegram-DB-1telegramrussian-telegram-chat-logs
russian-telegram-chat-logs
This dataset contains messages extracted from Telegram chat history, processed and ranked by their "semantic load" (information density).
Dataset Overview
The data is stored in Parquet format, which provides efficient storage and maintains data types. Each row represents a single message that has passed through several stages of cleaning and analysis.
Column
Type
Description
message
string
The cleaned text of the Telegram message… See the full description on the dataset page: https://huggingface.co/datasets/KvaytG/russian-telegram-chat-logs.telegram-spamtelegramtelegrampersian-telegram-corpus
PersianPoetryQuotes Dataset
Dataset Description
Dataset Summary
this is a collection of Persian messages extracted from various Telegram channels.
telegram_data_war_in_ukrainetelegramTelegram-DBTelegram-DB-1telegram-channel-dataset
Telegram AI Image Dataset — Cleaned for VLM LoRA Training
A cleaned dataset of 1,050 AI-generated images with their generation prompts, collected from a Chinese Telegram channel focused on GPT-Image-2 prompt engineering.
Each image is paired with a structured generation prompt in Chinese/English. Prompts have been cleaned of channel boilerplate — no bot instructions, hashtags, source credits, model name prefixes, emoji title lines, or channel footer ads.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GCStream/telegram-channel-dataset.telegram-chat-export-synthetic
Telegram Chat Export (Synthetic)
3,000 synthetic chat histories in the Telegram Desktop export (result.json) format, generated
for fine-tuning small language models on chat-style generation.
Used by: Offlin33er/SmolLM2-1.7B-telegram-chats (SFT with TRL) —
also available as GGUF quantizations.
Demo: Offlin33er/telegram-chat-generator renders chats live from that model.
Schema
Column
Type
Content
prompt
string
User turn asking for a chat history (topic… See the full description on the dataset page: https://huggingface.co/datasets/Offlin33er/telegram-chat-export-synthetic.Telegrampersian_arabic_pairs_telegramTelegram-DB
