CoolFace
Datasetpublic

kurumikz/telegram-corpus-russian-kazakh

๐Ÿ“ฆ Telegram Corpus KAZ_RU Telegram Corpus KAZ_RU is a raw multilingual dataset of Telegram messages in Russian and Kazakh, collected and assembled by Kurumikz. It contains over 1.4 million lines of informal, user-generated text extracted from a large .txt dump. The corpus is intended for experimentation in text generation, language modeling, and dialogue research. ๐Ÿ“Š Dataset Statistics Metric Value Description ๐Ÿ“„ Lines 1,493,124 Total number ofโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/telegram-corpus-russian-kazakh.

sourceHugging Facecc-by-nc-sa-4.0updated 1y agoView on Hugging Face
3likes67downloads
Dataset Card

๐Ÿ“ฆ Telegram Corpus KAZ_RU

Telegram Corpus KAZ_RU is a raw multilingual dataset of Telegram messages in Russian and Kazakh, collected and assembled by Kurumikz. It contains over 1.4 million lines of informal, user-generated text extracted from a large .txt dump. The corpus is intended for experimentation in text generation, language modeling, and dialogue research.


๐Ÿ“Š Dataset Statistics

MetricValueDescription
๐Ÿ“„ Lines1,493,124Total number of lines/messages
๐Ÿ“ฆ Characters33,753,922All characters including spaces and punctuation
๐Ÿ…ฐ๏ธ Letters (total)22,867,523All alphabetic characters
๐Ÿ”  Latin letters3,997,388English and other Latin-based letters
๐Ÿ”ก Cyrillic letters18,761,717Russian, Kazakh, and other Cyrillic letters
๐ŸŒ Other alphabets108,418Non-Latin/Cyrillic (e.g. Arabic, emoji, etc.)
๐Ÿ”ข Digits1,730,772All numeric characters
โฃ Spaces3,904,045Whitespace characters
โ€ข Periods456,884Sentence-ending punctuation
, Commas203,237Mid-sentence punctuation
๐Ÿงฉ Tokens (space-split)4,623,188Approximate token count using space delimiter

โš ๏ธ Notes & Warnings

This dataset is a raw prototype and has not been manually cleaned or annotated. It may contain:

  • โ€”Profanity, offensive language, and slang
  • โ€”Excessive whitespace and empty lines
  • โ€”Unstructured formatting, broken links, usernames, emojis
  • โ€”Mixed language usage (primarily Russian and Kazakh)
  • โ€”Repetitive or low-quality content

Use with caution for downstream tasks. Preprocessing is recommended before training or evaluation.


๐Ÿง  Intended Use

This corpus is suitable for:

  • โ€”Pretraining or fine-tuning LLMs (e.g. CompactLLM)
  • โ€”Studying informal multilingual text
  • โ€”Dialogue modeling and Telegram-style generation
  • โ€”Feature extraction and embedding training
  • โ€”Language modeling in Russian and Kazakh

๐Ÿ“„ License

This dataset is licensed under CC BY-NC-SA 4.0.

  • โ€”You must credit the author: Kurumikz
  • โ€”Commercial use is not allowed
  • โ€”Any derivative work must be shared under the same license

Full license text: Creative Commons Attribution-NonCommercial-ShareAlike 4.0