CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01esimijoq /Kazakh-Russian-Child-Directed-Speech-Corpus Kazakh-Russian Child-Directed Language Corpus Dataset Description The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization. The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts. This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.texttext-generation1M<n<10M0 likes1.1k downloads16d agoHugging Face02ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1k downloads1y agoHugging Face03RussianNLP /russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills - detection of natural language inference, commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating models and an overall leaderboard of transformer models for the Russian language.text-classification100K<n<1M35 likes792 downloads3y agoHugging Face04Den4ikAI /russian_dialogues_2 Den4ikAI/russian_dialogues_2 Датасет русских диалогов для обучения диалоговых моделей. Количество диалогов - 1.6 миллиона Формат датасета: { 'sample': ['Привет', 'Привет', 'Как дела?'] } Citation: @MISC{russian_instructions, author = {Denis Petrov}, title = {Russian context dialogues dataset for conversational agents}, url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2}, year = 2023 } texttext-generation1M<n<10M18 likes625 downloads2y agoHugging Face05wwewtech /russian-it-community-corpus 📦 Russian IT Community Corpus (RICC) Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture. The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.tabulartext-generation1M<n<10M1 likes339 downloads15d agoHugging Face06adeshkin /khakas-russian-parallel-corpus Khakas-Russian Parallel Corpus The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people. Dataset Overlap: The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.texttranslation100K<n<1M2 likes277 downloads13d agoHugging Face07adeshkin /khakas-russian-dict Khakas-Russian Dictionary (Dataset) Dataset Developer: Vasily AdeshkinContact for inquiries: adeshkin.vi@phystech.edu 📌 Important Notice & Citation When using this dataset, please be sure to cite this repository and the original dictionary: https://khakas.altaica.ru/dictionary/. Please note: Some optical character recognition (OCR) errors may still be present in the data. 🛠 Contribution & Authorship I am not the author of the original dictionary, but I have… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-dict.tabulartranslation10K<n<100K0 likes213 downloads5mo agoHugging Face08AITISPEC /physics-russian Оглавление Описание датасета Аннотация Ключевые особенности Статус перевода Методология перевода и верификации Ограничения и возможные погрешности Структура датасета Поля данных Использование Благодарности Лицензирование и авторские права Цитирование 📑 Оглавление Примеров Нажмите на тему, чтобы перейти к соответствующему разделу в полном отчёте. Квантовая механика Термодинамика Электромагнетизм Общая теория относительности Специальная теория относительности Атомная… See the full description on the dataset page: https://huggingface.co/datasets/AITISPEC/physics-russian.texttext-generation10K<n<100K0 likes207 downloads3mo agoHugging Face09drkolesnikov /russian-nmo-medical-mcq Russian NMO Medical MCQ Choose language / Выберите язык: Русский | English Русский Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа. В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA. Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.tabularquestion-answering1M<n<10M2 likes183 downloads4mo agoHugging Face10RafaelUI /russian_literature RusLit Corpus A corpus of Russian literature in clean text format. Description This dataset contains cleaned text (stripped of extraneous artifacts) collected from works by authors who passed away more than 70 years ago, placing them in the public domain. Some texts may contain semantic nonsense (e.g. OCR or digitization artifacts), as well as fragments of French, German, English, or Japanese text mixed in with the Russian. Structure Each record… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/russian_literature.texttext-generationn<1K1 likes172 downloads3mo agoHugging Face11issai /GSM8k_Kazakh_Russian Dataset Summary These are the machine-translated Kazakh and Russian versions of the GSM8K (Grade School Math 8K) dataset (test set). This dataset is used to test the mathematical reasoning of large language models in the Kazakh language. Specifically, GSM8K focuses on high-quality grade school math word problems that require multi-step reasoning to solve. Kazakh and Russian versions serve as a benchmark for evaluating how well models can perform mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/issai/GSM8k_Kazakh_Russian.texttext-generation1K<n<10K2 likes158 downloads1mo agoHugging Face12IgorVolochay /russian_jokestexttext-generation100K<n<1M13 likes153 downloads3y agoHugging Face13Lev384501 /russian_instructions_2_cleaned Russian Instructions Cleaned Очищенная версия Den4ikAI/russian_instructions_2. Что сделано Дедупликация по question (удалено ~45k) Удалены пустые question и answer Удалены ответы короче 100 и длиннее 4000 символов Конвертировано в chat-формат (messages: user/assistant) Статистика Метрика Значение Исходно 237 281 После чистки 138 973 Удалено 98 308 (41%) Формат JSONL, одна строка = один пример. {"messages":… See the full description on the dataset page: https://huggingface.co/datasets/Lev384501/russian_instructions_2_cleaned.texttext-generation100K<n<1M2 likes137 downloads6d agoHugging Face14eridai /russian_dpo_qa Format Each row contains: prompt chosen rejected Usage from datasets import load_dataset dataset = load_dataset("eridai/russian_dpo_qa") train = dataset["train"] texttext-generation1K<n<10K1 likes123 downloads13d agoHugging Face15attn-signs /russian-easy-instructions Easy Russian Instructions Fast instructions in conversational format, generaly used for injecting general knowledge.Dataset contains comprehensive instructions for easy tasks and general question answering Contents: Wikipedia / Factological knowledge History knowledge Basic programming understanding Basic math understanding Basic physics understanding Basic geography knowledge Basic biology knowledge Format: Formatted for instruction-tuning / instruction… See the full description on the dataset page: https://huggingface.co/datasets/attn-signs/russian-easy-instructions.textquestion-answering10K<n<100K1 likes108 downloads2y agoHugging Face16aglazkova /keyphrase_extraction_russianA dataset for extracting or generating keywords for scientific texts in Russian. It contains annotations of scientific articles from four domains: mathematics and computer science (a fragment of the Keyphrases CS&Math Russian dataset), history, linguistics, and medicine. For more details about the dataset please refer the original paper: https://arxiv.org/abs/2409.10640 Dataset Structure abstract, an abstract in a string format; keywords, a list of keyphrases provided by… See the full description on the dataset page: https://huggingface.co/datasets/aglazkova/keyphrase_extraction_russian.texttext-generation10K<n<100K5 likes98 downloads11mo agoHugging Face17samedad /mem-and-russian-jokes-dataset 2 июля 2025 Добавлено новых уникальных анекдотов: 1311963 Количество записей в датасете: 521904 Добавил датасет анекдотов от IgorVolochay/russian_jokes не понимаю как я прошел мимо него, там очень много шуток, сам Евгений Ваганович очень часто обращался к этому датасету... Но вот и мое время пришло. Плюс обновил старый датасет и разбавил формулировки новыми начальными вопросами human_prompts = [ "Расскажи шутку", "Расскажи анекдот", "Знаешь какой-нибудь прикол?", "Скажи что-нибудь смешное"… See the full description on the dataset page: https://huggingface.co/datasets/samedad/mem-and-russian-jokes-dataset.tabulartext-generation100K<n<1M4 likes90 downloads1y agoHugging Face18Kasymkhan /RussianFinancialNews RussianFinancialNews Датасет содержит 92,377 русскоязычных новостных статей на финансовую тематику, преимущественно про российский рынок ценных бумаг и российскую экономику. Набор данных может быть полезен для разных задач обработки естественного языка (NLP). A dataset containing 92,377 samples of Russian financial news articles. Each sample includes metadata and content fields that are useful for various Natural Language Processing (NLP) tasks. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/Kasymkhan/RussianFinancialNews.texttext-classification10K<n<100K0 likes77 downloads2y agoHugging Face19qwqeqw /Dataset_of_Russian_thinkingRu RTD Описание:Russian Thinking Dataset — это набор данных, предназначенный для обучения и тестирования моделей обработки естественного языка (NLP) на русском языке. Датасет ориентирован на задачи, связанные с генерацией текста, анализом диалогов и решением математических и логических задач. Основная информация: Сплит: train Количество записей: 147.046 Цели: Обучение моделей пониманию русского языка. Создание диалоговых систем с естественным взаимодействием.… See the full description on the dataset page: https://huggingface.co/datasets/qwqeqw/Dataset_of_Russian_thinking.texttext-generation100K<n<1M1 likes75 downloads10mo agoHugging Face20pavelfedortsov /russian-colloquial-sft-50k Russian colloquial SFT (50k, mat-free) English: ~50,000 chat SFT examples: rewrite formal Russian into casual Telegram-style Russian without profanity. Format Each line is JSON with a messages array (ShareGPT / TRL): { "messages": [ { "role": "user", "content": "Перепиши простым разговорным русским, как в переписке. Без мата и грубости. Сохрани смысл:\n<формальный текст>" }, { "role": "assistant", "content": "<разговорный ответ>"… See the full description on the dataset page: https://huggingface.co/datasets/pavelfedortsov/russian-colloquial-sft-50k.texttext-generation10K<n<100K0 likes74 downloads4mo agoHugging Face21kukunechka /russian-everyday-dialogues Russian Everyday Dialogues / Русские повседневные диалоги A dataset of 20 natural conversational Russian dialogues covering everyday situations. Создан носителем языка с 18-летним опытом работы в международной среде. Dataset Description This dataset contains short, natural Russian dialogues typical of everyday urban life in Russia. Each example reflects authentic spoken language patterns, not formal or literary Russian. Situations covered 🛒 Магазин / Shopping… See the full description on the dataset page: https://huggingface.co/datasets/kukunechka/russian-everyday-dialogues.text-generationn<1K0 likes73 downloads5mo agoHugging Face22Kemsekov /Corrupted-russian-word-documents-text-datasetThis is synthetic text-corruption dataset, based on Russian open official/government documents text. This dataset is intended to be used to train LLM to perform text-recovery task. All the errors in text is made solely in Russian sentences, hence ignoring any English sentence. Texts contains complex formatting, which is common for documents. Each line contains json object that have array messages value, which consists of role-based conversation. Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.texttext-generationn<1K2 likes63 downloads2y agoHugging Face23kurumikz /telegram-corpus-russian-kazakh 📦 Telegram Corpus KAZ_RU Telegram Corpus KAZ_RU is a raw multilingual dataset of Telegram messages in Russian and Kazakh, collected and assembled by Kurumikz. It contains over 1.4 million lines of informal, user-generated text extracted from a large .txt dump. The corpus is intended for experimentation in text generation, language modeling, and dialogue research. 📊 Dataset Statistics Metric Value Description 📄 Lines 1,493,124 Total number of lines/messages… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/telegram-corpus-russian-kazakh.texttext-generation1M<n<10M3 likes63 downloads11mo agoHugging Face24jebkoralav /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/jebkoralav/ru-big-russian-dataset.tabulartext-generation1M<n<10M0 likes62 downloads10mo agoHugging Face25kristaller486 /Nebo-T1-Russian Russian Description (English below) UPD: Dataset reuploaded, correct_format column added Nebo-T1-Russian (Вероятно) первый "longCoT" датасет для русского языка, созданный через Deeseek-R1 Подсказки взяты из датасета Sky-T1 и переведены через Llama3.3-70B Ответы и рассуждения сгенерированные Deeseek-R1 (685B) 16.4K сэмплов в целом, ≈12.4K только с русским языком (в остальных либо ответ, либо рассуждения на английском) Языки в ответе и рассуждениях размечены… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/Nebo-T1-Russian.texttext-generation10K<n<100K15 likes58 downloads2y agoHugging Face26evilfreelancer /MATH-500-Russian Карточка датасета MATH-500-Russian Перевод датасета HuggingFaceH4/MATH-500 на русский язык, был выполнен моделью qwen2.5:32b через скрипты EvilFreelancer/datasets-translator. Данный набор данных содержит подмножество из 500 задач из теста MATH, который OpenAI создал для статьи Let's Verify Step by Step и переведённых на русский язык. Подробности в их репозиторий на GitHub. texttext-generationn<1K3 likes56 downloads2y agoHugging Face27ScoutieAutoML /russian_oil_gas_news_telegram_dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.tabulartext-classification10K<n<100K7 likes53 downloads2y agoHugging Face28opendatalab /WanJuan-Russian 💡 Introduction WanJuan-Russian(万卷丝路-俄语)corpus, with a volume exceeding 280GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Russian.text-generation6 likes50 downloads1y agoHugging Face29ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes46 downloads2y agoHugging Face30Embim /Russian_spell_dataset Russian Spell Dataset Датасет для обучения и дообучения моделей исправления ошибок автоматического распознавания русской речи. Содержание Датасет содержит 1 139 пар текстов в формате JSONL: input — текст с опечатками, ошибками распознавания речи, пропущенной пунктуацией или неправильным регистром; output — исправленная версия текста. Также есть примеры без ошибок, где input совпадает с output. Это помогает модели не вносить исправления в уже корректный текст.… See the full description on the dataset page: https://huggingface.co/datasets/Embim/Russian_spell_dataset.texttext-generation1K<n<10K0 likes46 downloads7d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.