CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RussianNLP /coat Dataset Card for CoAT🧥 Dataset Description CoAT🧥 (Corpus of Artificial Texts) is a large-scale corpus for Russian, which consists of 246k human-written texts from publicly available resources and artificial texts generated by 13 neural models, varying in the number of parameters, architecture choices, pre-training objectives, and downstream applications. Each model is fine-tuned for one or more of six natural language generation tasks, ranging from paraphrase generation… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/coat.tabulartext-classification100K<n<1M4 likes2.2k downloads10mo agoHugging Face02ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1k downloads1y agoHugging Face03RussianNLP /rublimp RuBLiMP Dataset Description RuBLiMP, or Russian Benchmark of Linguistic Minimal Pairs, is the first diverse and large-scale benchmark of minimal pairs in Russian. RuBLiMP includes 45k minimal pairs of sentences that differ in grammaticality and isolate morphological, syntactic, or semantic phenomena. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/rublimp.tabular10K<n<100K4 likes727 downloads1y agoHugging Face04rustemgareev /russian-names Russian Names with Popularity Scores Description This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025. Usage The dataset can be loaded using the Hugging Face datasets library. from datasets import load_dataset dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.tabularother10K<n<100K0 likes410 downloads1y agoHugging Face05wwewtech /russian-it-community-corpus 📦 Russian IT Community Corpus (RICC) Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture. The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.tabulartext-generation1M<n<10M1 likes331 downloads16d agoHugging Face06Alexey5676 /russian-supreme-court-plenum-acts Plenum Resolutions of the Supreme Court of Russia (1961–2026) Every act published in the «Постановления Пленума» section of the Russian Supreme Court's website: 1,504 records — 1,503 plenum resolutions plus 1 meeting agenda — with full texts, metadata and the court's original attachments. Coverage 1961–2026; completeness verified against the court's own index at collection time (the section reported exactly 1,504 documents). Постановления Пленума ВС РФ — руководящие разъяснения… See the full description on the dataset page: https://huggingface.co/datasets/Alexey5676/russian-supreme-court-plenum-acts.documentsummarization1K<n<10K2 likes309 downloads8d agoHugging Face07PleIAs /Russian-PD 🇷🇺 Russian Public Domain 🇷🇺 Russian-Public Domain or Russian-PD is a large collection aiming to aggregate all Russian monographies and periodicals in the public domain. Dataset summary The collection contains 8525 titles making up 995,163,165 words recovered from the Internet Archive. Each parquet file has the full text of 2,000 books selected at random. Curation method The composition of the dataset adheres to the criteria for public domain works in the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Russian-PD.tabular1K<n<10K6 likes289 downloads2y agoHugging Face08OVHaiLLM /russian_super_glue RussianNLP/russian_super_glue (script-less mirror) Script-less mirror of the source dataset. All original columns preserved. All original splits preserved (no renaming). Each config is a subset. tabular100K<n<1M0 likes236 downloads1y agoHugging Face09bethrezen /ru-big-russian-dataset-16k-tokens-limittabular1M<n<10M0 likes214 downloads1y agoHugging Face10adeshkin /khakas-russian-dict Khakas-Russian Dictionary (Dataset) Dataset Developer: Vasily AdeshkinContact for inquiries: adeshkin.vi@phystech.edu 📌 Important Notice & Citation When using this dataset, please be sure to cite this repository and the original dictionary: https://khakas.altaica.ru/dictionary/. Please note: Some optical character recognition (OCR) errors may still be present in the data. 🛠 Contribution & Authorship I am not the author of the original dictionary, but I have… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-dict.tabulartranslation10K<n<100K0 likes214 downloads5mo agoHugging Face11drkolesnikov /russian-nmo-medical-mcq Russian NMO Medical MCQ Choose language / Выберите язык: Русский | English Русский Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа. В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA. Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.tabularquestion-answering1M<n<10M2 likes174 downloads4mo agoHugging Face12ScoutieAutoML /russian-news-telegram-dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic News and Media, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the channel. views… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian-news-telegram-dataset.tabulartext-classification10K<n<100K6 likes164 downloads2y agoHugging Face13ZeroAgency /ru-big-russian-dataset-v1.1tabular1M<n<10M5 likes155 downloads1y agoHugging Face14justicedao /ipfs_russia_laws_ir Russia legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_russia_laws (revision c81227eb09f7baebbad861f65bcf5d2a92408602) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Russia prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_russia_laws_ir.tabulartext-retrieval100K<n<1M0 likes124 downloads2d agoHugging Face15Roman-Kpro /russian-business-registries Russian Business Registries — Aggregated Statistics Aggregated, ready-to-analyse slices of Russian state registers. Every figure comes from an official open-data source; nothing here is modelled, imputed or estimated. Individual companies are not published — only aggregates, with one deliberate exception described below. Собрано из открытых данных российских госреестров. Все цифры — из официальных источников, без моделирования и досчётов. Публикуются агрегаты, не сведения об… See the full description on the dataset page: https://huggingface.co/datasets/Roman-Kpro/russian-business-registries.tabulartabular-classification10K<n<100K0 likes116 downloads11d agoHugging Face16IvanFed /russian-toxic-comments-multilabel Russian Toxic Comments Multi-label Dataset Dataset Description Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности. Цель Обучение модели для автоматического обнаружения трех типов токсичного контента: Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.tabulartext-classification100K<n<1M1 likes114 downloads2mo agoHugging Face17Romjiik /Russian_bank_reviews Dataset Card for bank reviews dataset Dataset Summary The dataset is collected from the banki.ru website. It contains customer reviews of various banks. In total, the dataset contains 12399 reviews. The dataset is suitable for sentiment classification. The dataset contains this fields - bank name, username, review title, review text, review time, number of views, number of comments, review rating set by the user, as well as ratings for special categories… See the full description on the dataset page: https://huggingface.co/datasets/Romjiik/Russian_bank_reviews.tabulartext-classification10K<n<100K7 likes111 downloads3y agoHugging Face18Ramilles /russia_housingtabular10M<n<100M0 likes96 downloads10d agoHugging Face19samedad /mem-and-russian-jokes-dataset 2 июля 2025 Добавлено новых уникальных анекдотов: 1311963 Количество записей в датасете: 521904 Добавил датасет анекдотов от IgorVolochay/russian_jokes не понимаю как я прошел мимо него, там очень много шуток, сам Евгений Ваганович очень часто обращался к этому датасету... Но вот и мое время пришло. Плюс обновил старый датасет и разбавил формулировки новыми начальными вопросами human_prompts = [ "Расскажи шутку", "Расскажи анекдот", "Знаешь какой-нибудь прикол?", "Скажи что-нибудь смешное"… See the full description on the dataset page: https://huggingface.co/datasets/samedad/mem-and-russian-jokes-dataset.tabulartext-generation100K<n<1M4 likes89 downloads1y agoHugging Face20nevmenandr /russian-20th-century-bigrams Русскоязычные биграммы XX века Общие замечания В этом датасете содержатся преобразованные в более подходящий для исследования вид биграммы на русском языке и их частотности из коллекции Google Ngrams с 1918 до 2010 года. Такой выбор обусловлен диапазоном дат, в рамках которого в русском языке соблюдается современный орфографический режим. Дореформенная орфография хуже распознается системами OCR, с ней невозможно работать as is современными средствами обработки… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-20th-century-bigrams.tabular10M<n<100M1 likes70 downloads11mo agoHugging Face21ru-dataset /dzen-russian-articles Dzen Russian Articles Dataset Русскоязычные статьи с dzen.ru. Датасет в активном сборе — новые статьи добавляются регулярно, объём постоянно растёт. Как устроен парсинг Статьи собираются с dzen.ru и обрабатываются через Gemini в качестве движка извлечения: модель очищает текст, разбивает на абзацы, определяет категорию, теги, ключевые слова и тональность. Поле extractor содержит название используемой модели (gemini/gemini-3.5-flash-lite). Структура… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/dzen-russian-articles.tabulartext-classificationn<1K1 likes69 downloads2mo agoHugging Face22dbrovkin /toxic-russian-comments-multilabel Russian Toxic Comments Multi-Label Balanced Dataset Описание Этот датасет создан для задачи Multi-Task классификации токсичности русскоязычных комментариев. Датасет содержит сбалансированные примеры с тремя бинарными метками: profanity: наличие нецензурной лексики (мат) threat: наличие угроз illegal: запросы на незаконные действия (прокси-метка на основе THREAT + INSULT) Структура данных Датасет содержит следующие поля: Поле Тип Описание… See the full description on the dataset page: https://huggingface.co/datasets/dbrovkin/toxic-russian-comments-multilabel.tabulartext-classification100K<n<1M1 likes68 downloads2mo agoHugging Face23ScoutieAutoML /russian_oil_gas_news_telegram_dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.tabulartext-classification10K<n<100K7 likes62 downloads2y agoHugging Face24jebkoralav /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/jebkoralav/ru-big-russian-dataset.tabulartext-generation1M<n<10M0 likes60 downloads10mo agoHugging Face25Vikhrmodels /russian-asr-leaderboardtabularn<1K0 likes59 downloads1y agoHugging Face26issai /MMLU-Pro_Kazakh_Russian Dataset Summary These are the machine-translated Kazakh and Russian versions of the MMLU-Pro (Massive Multitask Language Understanding Pro) dataset (test set). These datasets are used to test the world knowledge and problem-solving capabilities of large language models across a vast range of subjects in the Kazakh and Russian languages. As an enhanced version of the original MMLU, it serves as a more rigorous benchmark for evaluating how well models understand complex academic… See the full description on the dataset page: https://huggingface.co/datasets/issai/MMLU-Pro_Kazakh_Russian.tabularquestion-answering10K<n<100K2 likes58 downloads1mo agoHugging Face27k-mktr /russian_synodal_ru Russian Synodal Bible (1876) Description The Russian Synodal Bible (Синодальный перевод) is the official translation of the Bible into Russian, authorized by the Holy Synod of the Russian Orthodox Church. The translation was completed in 1876, with the New Testament published earlier in 1862. It was translated from the original Hebrew, Aramaic, and Greek texts by a team of scholars from the Russian Orthodox Church and theological academies. The Synodal Bible… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/russian_synodal_ru.tabular10K<n<100K0 likes58 downloads2mo agoHugging Face28daniilakk /Russia_Real_Estate_2018_2021 Context The dataset consists of lists of unique objects of popular portals for the sale of real estate in Russia. More than 540 thousand objects. The dataset contains 540000 real estate objects in Russia. Content The Russian real estate market has a relatively short history. In the Soviet era, all properties were state-owned; people only had the right to use them with apartments allocated based on one's place of work. As a result, options for moving were fairly limited.… See the full description on the dataset page: https://huggingface.co/datasets/daniilakk/Russia_Real_Estate_2018_2021.tabular1M<n<10M3 likes55 downloads4y agoHugging Face29DBQ /Louis.Vuitton.Product.prices.Russia Louis Vuitton web scraped data About the website The luxury fashion industry in the EMEA region, particularly in Russia, is characterized by a growing demand for high-end products from renowned brands. Louis Vuitton, a global leader in this industry, caters to this escalating demand through their extensive range of luxury clothing, accessories, and luggage. The brand has significantly increased its presence in Russia by leveraging the power of Ecommerce, effectively… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Louis.Vuitton.Product.prices.Russia.imagetext-classification1K<n<10K2 likes51 downloads3y agoHugging Face30kaengreg /rus-scifact-qrelstabular1K<n<10K0 likes48 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.