CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01esimijoq /Kazakh-Russian-Child-Directed-Speech-Corpus Kazakh-Russian Child-Directed Language Corpus Dataset Description The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization. The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts. This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.texttext-generation1M<n<10M0 likes1.1k downloads17d agoHugging Face02ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1k downloads1y agoHugging Face03Den4ikAI /russian_dialogues_2 Den4ikAI/russian_dialogues_2 Датасет русских диалогов для обучения диалоговых моделей. Количество диалогов - 1.6 миллиона Формат датасета: { 'sample': ['Привет', 'Привет', 'Как дела?'] } Citation: @MISC{russian_instructions, author = {Denis Petrov}, title = {Russian context dialogues dataset for conversational agents}, url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2}, year = 2023 } texttext-generation1M<n<10M18 likes610 downloads2y agoHugging Face04wwewtech /russian-it-community-corpus 📦 Russian IT Community Corpus (RICC) Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture. The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.tabulartext-generation1M<n<10M1 likes331 downloads17d agoHugging Face05adeshkin /khakas-russian-parallel-corpus Khakas-Russian Parallel Corpus The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people. Dataset Overlap: The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.texttranslation100K<n<1M2 likes275 downloads15d agoHugging Face06adeshkin /khakas-russian-dict Khakas-Russian Dictionary (Dataset) Dataset Developer: Vasily AdeshkinContact for inquiries: adeshkin.vi@phystech.edu 📌 Important Notice & Citation When using this dataset, please be sure to cite this repository and the original dictionary: https://khakas.altaica.ru/dictionary/. Please note: Some optical character recognition (OCR) errors may still be present in the data. 🛠 Contribution & Authorship I am not the author of the original dictionary, but I have… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-dict.tabulartranslation10K<n<100K0 likes214 downloads5mo agoHugging Face07RafaelUI /russian_literature RusLit Corpus A corpus of Russian literature in clean text format. Description This dataset contains cleaned text (stripped of extraneous artifacts) collected from works by authors who passed away more than 70 years ago, placing them in the public domain. Some texts may contain semantic nonsense (e.g. OCR or digitization artifacts), as well as fragments of French, German, English, or Japanese text mixed in with the Russian. Structure Each record… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/russian_literature.texttext-generationn<1K1 likes209 downloads3mo agoHugging Face08AITISPEC /physics-russian Оглавление Описание датасета Аннотация Ключевые особенности Статус перевода Методология перевода и верификации Ограничения и возможные погрешности Структура датасета Поля данных Использование Благодарности Лицензирование и авторские права Цитирование 📑 Оглавление Примеров Нажмите на тему, чтобы перейти к соответствующему разделу в полном отчёте. Квантовая механика Термодинамика Электромагнетизм Общая теория относительности Специальная теория относительности Атомная… See the full description on the dataset page: https://huggingface.co/datasets/AITISPEC/physics-russian.texttext-generation10K<n<100K0 likes208 downloads3mo agoHugging Face09drkolesnikov /russian-nmo-medical-mcq Russian NMO Medical MCQ Choose language / Выберите язык: Русский | English Русский Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа. В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA. Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.tabularquestion-answering1M<n<10M2 likes174 downloads4mo agoHugging Face10issai /GSM8k_Kazakh_Russian Dataset Summary These are the machine-translated Kazakh and Russian versions of the GSM8K (Grade School Math 8K) dataset (test set). This dataset is used to test the mathematical reasoning of large language models in the Kazakh language. Specifically, GSM8K focuses on high-quality grade school math word problems that require multi-step reasoning to solve. Kazakh and Russian versions serve as a benchmark for evaluating how well models can perform mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/issai/GSM8k_Kazakh_Russian.texttext-generation1K<n<10K2 likes161 downloads1mo agoHugging Face11IgorVolochay /russian_jokestexttext-generation100K<n<1M13 likes158 downloads3y agoHugging Face12Lev384501 /russian_instructions_2_cleaned Russian Instructions Cleaned Очищенная версия Den4ikAI/russian_instructions_2. Что сделано Дедупликация по question (удалено ~45k) Удалены пустые question и answer Удалены ответы короче 100 и длиннее 4000 символов Конвертировано в chat-формат (messages: user/assistant) Статистика Метрика Значение Исходно 237 281 После чистки 138 973 Удалено 98 308 (41%) Формат JSONL, одна строка = один пример. {"messages":… See the full description on the dataset page: https://huggingface.co/datasets/Lev384501/russian_instructions_2_cleaned.texttext-generation100K<n<1M2 likes149 downloads7d agoHugging Face13eridai /russian_dpo_qa Format Each row contains: prompt chosen rejected Usage from datasets import load_dataset dataset = load_dataset("eridai/russian_dpo_qa") train = dataset["train"] texttext-generation1K<n<10K1 likes122 downloads14d agoHugging Face14attn-signs /russian-easy-instructions Easy Russian Instructions Fast instructions in conversational format, generaly used for injecting general knowledge.Dataset contains comprehensive instructions for easy tasks and general question answering Contents: Wikipedia / Factological knowledge History knowledge Basic programming understanding Basic math understanding Basic physics understanding Basic geography knowledge Basic biology knowledge Format: Formatted for instruction-tuning / instruction… See the full description on the dataset page: https://huggingface.co/datasets/attn-signs/russian-easy-instructions.textquestion-answering10K<n<100K1 likes111 downloads2y agoHugging Face15aglazkova /keyphrase_extraction_russianA dataset for extracting or generating keywords for scientific texts in Russian. It contains annotations of scientific articles from four domains: mathematics and computer science (a fragment of the Keyphrases CS&Math Russian dataset), history, linguistics, and medicine. For more details about the dataset please refer the original paper: https://arxiv.org/abs/2409.10640 Dataset Structure abstract, an abstract in a string format; keywords, a list of keyphrases provided by… See the full description on the dataset page: https://huggingface.co/datasets/aglazkova/keyphrase_extraction_russian.texttext-generation10K<n<100K5 likes104 downloads11mo agoHugging Face16samedad /mem-and-russian-jokes-dataset 2 июля 2025 Добавлено новых уникальных анекдотов: 1311963 Количество записей в датасете: 521904 Добавил датасет анекдотов от IgorVolochay/russian_jokes не понимаю как я прошел мимо него, там очень много шуток, сам Евгений Ваганович очень часто обращался к этому датасету... Но вот и мое время пришло. Плюс обновил старый датасет и разбавил формулировки новыми начальными вопросами human_prompts = [ "Расскажи шутку", "Расскажи анекдот", "Знаешь какой-нибудь прикол?", "Скажи что-нибудь смешное"… See the full description on the dataset page: https://huggingface.co/datasets/samedad/mem-and-russian-jokes-dataset.tabulartext-generation100K<n<1M4 likes89 downloads1y agoHugging Face17qwqeqw /Dataset_of_Russian_thinkingRu RTD Описание:Russian Thinking Dataset — это набор данных, предназначенный для обучения и тестирования моделей обработки естественного языка (NLP) на русском языке. Датасет ориентирован на задачи, связанные с генерацией текста, анализом диалогов и решением математических и логических задач. Основная информация: Сплит: train Количество записей: 147.046 Цели: Обучение моделей пониманию русского языка. Создание диалоговых систем с естественным взаимодействием.… See the full description on the dataset page: https://huggingface.co/datasets/qwqeqw/Dataset_of_Russian_thinking.texttext-generation100K<n<1M1 likes83 downloads10mo agoHugging Face18Kasymkhan /RussianFinancialNews RussianFinancialNews Датасет содержит 92,377 русскоязычных новостных статей на финансовую тематику, преимущественно про российский рынок ценных бумаг и российскую экономику. Набор данных может быть полезен для разных задач обработки естественного языка (NLP). A dataset containing 92,377 samples of Russian financial news articles. Each sample includes metadata and content fields that are useful for various Natural Language Processing (NLP) tasks. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/Kasymkhan/RussianFinancialNews.texttext-classification10K<n<100K0 likes80 downloads2y agoHugging Face19pavelfedortsov /russian-colloquial-sft-50k Russian colloquial SFT (50k, mat-free) English: ~50,000 chat SFT examples: rewrite formal Russian into casual Telegram-style Russian without profanity. Format Each line is JSON with a messages array (ShareGPT / TRL): { "messages": [ { "role": "user", "content": "Перепиши простым разговорным русским, как в переписке. Без мата и грубости. Сохрани смысл:\n<формальный текст>" }, { "role": "assistant", "content": "<разговорный ответ>"… See the full description on the dataset page: https://huggingface.co/datasets/pavelfedortsov/russian-colloquial-sft-50k.texttext-generation10K<n<100K0 likes72 downloads4mo agoHugging Face20Kemsekov /Corrupted-russian-word-documents-text-datasetThis is synthetic text-corruption dataset, based on Russian open official/government documents text. This dataset is intended to be used to train LLM to perform text-recovery task. All the errors in text is made solely in Russian sentences, hence ignoring any English sentence. Texts contains complex formatting, which is common for documents. Each line contains json object that have array messages value, which consists of role-based conversation. Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.texttext-generationn<1K2 likes65 downloads2y agoHugging Face21kurumikz /telegram-corpus-russian-kazakh 📦 Telegram Corpus KAZ_RU Telegram Corpus KAZ_RU is a raw multilingual dataset of Telegram messages in Russian and Kazakh, collected and assembled by Kurumikz. It contains over 1.4 million lines of informal, user-generated text extracted from a large .txt dump. The corpus is intended for experimentation in text generation, language modeling, and dialogue research. 📊 Dataset Statistics Metric Value Description 📄 Lines 1,493,124 Total number of lines/messages… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/telegram-corpus-russian-kazakh.texttext-generation1M<n<10M3 likes63 downloads11mo agoHugging Face22ScoutieAutoML /russian_oil_gas_news_telegram_dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.tabulartext-classification10K<n<100K7 likes62 downloads2y agoHugging Face23jebkoralav /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/jebkoralav/ru-big-russian-dataset.tabulartext-generation1M<n<10M0 likes60 downloads10mo agoHugging Face24kristaller486 /Nebo-T1-Russian Russian Description (English below) UPD: Dataset reuploaded, correct_format column added Nebo-T1-Russian (Вероятно) первый "longCoT" датасет для русского языка, созданный через Deeseek-R1 Подсказки взяты из датасета Sky-T1 и переведены через Llama3.3-70B Ответы и рассуждения сгенерированные Deeseek-R1 (685B) 16.4K сэмплов в целом, ≈12.4K только с русским языком (в остальных либо ответ, либо рассуждения на английском) Языки в ответе и рассуждениях размечены… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/Nebo-T1-Russian.texttext-generation10K<n<100K15 likes58 downloads2y agoHugging Face25evilfreelancer /MATH-500-Russian Карточка датасета MATH-500-Russian Перевод датасета HuggingFaceH4/MATH-500 на русский язык, был выполнен моделью qwen2.5:32b через скрипты EvilFreelancer/datasets-translator. Данный набор данных содержит подмножество из 500 задач из теста MATH, который OpenAI создал для статьи Let's Verify Step by Step и переведённых на русский язык. Подробности в их репозиторий на GitHub. texttext-generationn<1K3 likes56 downloads2y agoHugging Face26ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes48 downloads2y agoHugging Face27Embim /Russian_spell_dataset Russian Spell Dataset Датасет для обучения и дообучения моделей исправления ошибок автоматического распознавания русской речи. Содержание Датасет содержит 1 139 пар текстов в формате JSONL: input — текст с опечатками, ошибками распознавания речи, пропущенной пунктуацией или неправильным регистром; output — исправленная версия текста. Также есть примеры без ошибок, где input совпадает с output. Это помогает модели не вносить исправления в уже корректный текст.… See the full description on the dataset page: https://huggingface.co/datasets/Embim/Russian_spell_dataset.texttext-generation1K<n<10K0 likes47 downloads8d agoHugging Face28DaniilMako /even-russian-instructions Even Russian Instructions Инструкционный датасет для fine-tuning VLM-моделей (Qwen3-VL) на эвенском языке — критически исчезающем тунгусо-маньчжурском языке Сибири.Примеры сгенерировано путем синтеза шаблонным методом на основе параллельного русско-эвенского корпуса. Структура датасета Три предварительно разделённых сплита для curriculum learning: Сплит Примеров Размер train 96 327 2.6 MB val 5 665 175 KB test 11 331 341 KB Поля… See the full description on the dataset page: https://huggingface.co/datasets/DaniilMako/even-russian-instructions.texttranslation100K<n<1M0 likes42 downloads2mo agoHugging Face29data-is-better-together /MPEP_RUSSIAN Dataset Card for MPEP_RUSSIAN This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/MPEP_RUSSIAN.texttext-generationn<1K4 likes41 downloads2y agoHugging Face30issai /RealWorldQA_Kazakh_Russian Dataset Summary These are the machine-translated Kazakh and Russian versions of the RealWorldQA dataset. As a multimodal benchmark, RealWorldQA is specifically designed to evaluate the real-world spatial understanding and visual reasoning of models. Kazakh and Russian versions serve as a benchmark for evaluating how well models can understand physical environments, spatial relationships, and object attributes based on real-world images when prompted in Kazakh or Russian.… See the full description on the dataset page: https://huggingface.co/datasets/issai/RealWorldQA_Kazakh_Russian.imagequestion-answering1K<n<10K0 likes35 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.