kurumikz/telegram-corpus-russian-kazakh
๐ฆ Telegram Corpus KAZ_RU Telegram Corpus KAZ_RU is a raw multilingual dataset of Telegram messages in Russian and Kazakh, collected and assembled by Kurumikz. It contains over 1.4 million lines of informal, user-generated text extracted from a large .txt dump. The corpus is intended for experimentation in text generation, language modeling, and dialogue research. ๐ Dataset Statistics Metric Value Description ๐ Lines 1,493,124 Total number ofโฆ See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/telegram-corpus-russian-kazakh.
๐ฆ Telegram Corpus KAZ_RU
Telegram Corpus KAZ_RU is a raw multilingual dataset of Telegram messages in Russian and Kazakh, collected and assembled by Kurumikz. It contains over 1.4 million lines of informal, user-generated text extracted from a large .txt dump. The corpus is intended for experimentation in text generation, language modeling, and dialogue research.
๐ Dataset Statistics
โ ๏ธ Notes & Warnings
This dataset is a raw prototype and has not been manually cleaned or annotated. It may contain:
- Profanity, offensive language, and slang
- Excessive whitespace and empty lines
- Unstructured formatting, broken links, usernames, emojis
- Mixed language usage (primarily Russian and Kazakh)
- Repetitive or low-quality content
Use with caution for downstream tasks. Preprocessing is recommended before training or evaluation.
๐ง Intended Use
This corpus is suitable for:
- Pretraining or fine-tuning LLMs (e.g. CompactLLM)
- Studying informal multilingual text
- Dialogue modeling and Telegram-style generation
- Feature extraction and embedding training
- Language modeling in Russian and Kazakh
๐ License
This dataset is licensed under CC BY-NC-SA 4.0.
- You must credit the author: Kurumikz
- Commercial use is not allowed
- Any derivative work must be shared under the same license
Full license text: Creative Commons Attribution-NonCommercial-ShareAlike 4.0
