datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
russian-names
Russian Names with Popularity Scores
Description
This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.russian-business-registries
Russian Business Registries — Aggregated Statistics
Aggregated, ready-to-analyse slices of Russian state registers. Every figure
comes from an official open-data source; nothing here is modelled, imputed or
estimated. Individual companies are not published — only aggregates, with one
deliberate exception described below.
Собрано из открытых данных российских госреестров. Все цифры — из официальных
источников, без моделирования и досчётов. Публикуются агрегаты, не сведения
об… See the full description on the dataset page: https://huggingface.co/datasets/Roman-Kpro/russian-business-registries.Russian_bank_reviews
Dataset Card for bank reviews dataset
Dataset Summary
The dataset is collected from the banki.ru website.
It contains customer reviews of various banks. In total, the dataset contains 12399 reviews.
The dataset is suitable for sentiment classification.
The dataset contains this fields - bank name, username, review title, review text, review time, number of views,
number of comments, review rating set by the user, as well as ratings for special categories… See the full description on the dataset page: https://huggingface.co/datasets/Romjiik/Russian_bank_reviews.russian-20th-century-bigrams
Русскоязычные биграммы XX века
Общие замечания
В этом датасете содержатся преобразованные в более подходящий для исследования вид биграммы на русском языке и их частотности из коллекции Google Ngrams с 1918 до 2010 года. Такой выбор обусловлен диапазоном дат, в рамках которого в русском языке соблюдается современный орфографический режим. Дореформенная орфография хуже распознается системами OCR, с ней невозможно работать as is современными средствами обработки… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-20th-century-bigrams.russian_oil_gas_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.Russia_Real_Estate_2018_2021
Context
The dataset consists of lists of unique objects of popular portals for the sale of real estate in Russia. More than 540 thousand objects.
The dataset contains 540000 real estate objects in Russia.
Content
The Russian real estate market has a relatively short history. In the Soviet era, all properties were state-owned; people only had the right to use them with apartments allocated based on one's place of work. As a result, options for moving were fairly limited.… See the full description on the dataset page: https://huggingface.co/datasets/daniilakk/Russia_Real_Estate_2018_2021.rus-scifact-qrelsscoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.Russian_Sensitive_Topics
General concept of the model
Sensitive topics are such topics that have a high chance of initiating a toxic conversation: homophobia, politics, racism, etc. This dataset uses 18 topics.
More details can be found in this article presented at the workshop for Balto-Slavic NLP at the EACL-2021 conference.
This paper presents the first version of this dataset. Here you can see the last version of the dataset which is significantly larger and also properly filtered.… See the full description on the dataset page: https://huggingface.co/datasets/NiGuLa/Russian_Sensitive_Topics.russian-names
Russian Names with Popularity Scores
Description
This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-names"… See the full description on the dataset page: https://huggingface.co/datasets/ssuverin/russian-names.Russian_Romantic_Dialogue_Dataset
💬 Russian Romantic Dialogue Dataset — Real AI-to-Human Conversations
🧩 Dataset Summary
This dataset contains annotated Russian-language romantic dialogues produced by a proprietary AI dating assistant operating in production across multiple platforms. Each dialogue is a real interaction between an AI system and a real woman responding in natural conditions.
The women respond naturally, producing authentic emotional dynamics, trust signals, resistance patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/DenSeduct/Russian_Romantic_Dialogue_Dataset.bible-lezghian-russianrussian-given-names-nen
Russian Given Names (NEN) — 1,551 names with meanings and ZAGS popularity
1,551 Russian given names (809 male, 742 female) with origin, short meaning, an editorial etymology note, diminutive and international forms, and popularity ranks among newborns based on open data from Moscow civil registry offices (ZAGS).
Curated by the editorial team of NEN («Нет, это нормально»), a Russian parenting magazine. Every record links to a full name page at n-e-n.ru/imena — the living catalog… See the full description on the dataset page: https://huggingface.co/datasets/MentalTech/russian-given-names-nen.russian_jokes_with_vectors
Description in English:
The dataset is collected from Russian-language Telegram channels with jokes and anecdotes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_jokes_with_vectors.benchmark-2-russian-m2mInfo:
Translated on Russian by facebook/m2m100_418M model
Source: xTRam1/safe-guard-prompt-injection
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Russian translated by facebook/m2m100_418M
score_ru_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-2-russian-m2m.russian_events_vectors
Description in English:
The dataset is collected from Russian-language Telegram channels with information about various events and happenings in Russian regions,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_events_vectors.tyvan-russian-parallel-50kA 50K sample from the Russian-Tyvan parallel corpus collected at https://tyvan.ru.
Russia_Real_Estate_2021Real estate ads in Russia are published on the websites avito.ru, realty.yandex.ru, cian.ru, sob.ru, youla.ru, n1.ru, moyareklama.ru. The ads-api.ru service allows you to upload real estate ads for a fee. The parser of the service works strangely and duplicates real estate ads in the database if the authors extended them after some time. Also in the Russian market there are a lot of outbids (bad realtors) who steal ads and publish them on their own behalf. Before publishing this dataset, my… See the full description on the dataset page: https://huggingface.co/datasets/daniilakk/Russia_Real_Estate_2021.scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.Russian_apartment_price
How to Use the Russian Apartment Price Dataset (Kvartis)
Dataset: main_data.csvGitHub Repository: zect-project/top_datasetsSize: 4,000 real estate listings (2025–2026 data)Target variable: real_price (price in Russian Rubles, ₽)
📋 Dataset Overview
This dataset contains detailed information about apartments for sale in five major Russian cities:
Moscow
Petersburg (St. Petersburg)
Novosibirsk
Yekaterinburg
Kazan
It is perfect for:
Price prediction… See the full description on the dataset page: https://huggingface.co/datasets/zect7/Russian_apartment_price.EyeWino
EyeWino
EyeWino is a new dataset based on the data from human eye-tracking for anaphora resolution.
Dataset Description
The Russian Winograd Schema Challenge dataset from TAPE (Taktasheva et al., 2022) was utilized for the anaphora resolution task to gather information on participants' eye movements.
The final dataset consists of 296 sentence-question pairs, which contain 9319 words and 148 unique sentences. The average number of participants per word is 48. The total… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/EyeWino.Russian2Evenkirussian_englishscoutieDataset_english_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.benchmark-1-russian-m2mInfo:
Translated on Russian by facebook/m2m100_418M model
Source: jayavibhav/prompt-injection-safety
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Russian translated by facebook/m2m100_418M
score_ru_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-russian-m2m.benchmark-4-russian-gtInfo:
Translated on Russian by Google Translate
Source: nvidia/Aegis-AI-Content-Safety-Dataset-2.0
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 1,000 prompts (500 safe / 500… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-4-russian-gt.benchmark-4-russian-m2mInfo:
Translated on Russian by facebook/m2m100_418M model
Source: nvidia/Aegis-AI-Content-Safety-Dataset-2.0
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 1,000 prompts (500… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-4-russian-m2m.Russian_bank_reviews
Dataset Card for bank reviews dataset
Dataset Summary
The dataset is collected from the banki.ru website.
It contains customer reviews of various banks. In total, the dataset contains 12399 reviews.
The dataset is suitable for sentiment classification.
The dataset contains this fields - bank name, username, review title, review text, review time, number of views,
number of comments, review rating set by the user, as well as ratings for special categories… See the full description on the dataset page: https://huggingface.co/datasets/ayasnoblin/Russian_bank_reviews.russian_trolls_preproc
Dataset Card for Dataset Name
The Russian Trolls twitter dataset as released and reported by NBC News.
From the original data file header:
"Tweets from confirmed Russian trolls, shows only username, timestamp (in UTC), tweet text, and number of times tweet was retweeted and favorited according to our data",,,,,,,,,,,,,,,,,
From NBC News' story: https://www.nbcnews.com/tech/social-media/now-available-more-200-000-deleted-russian-troll-tweets-n844731,,,,,,,,,,,,,,,,,
"If you… See the full description on the dataset page: https://huggingface.co/datasets/Kristijan/russian_trolls_preproc.benchmark-1-russian-gtInfo:
Translated on Russian by Google Translate
Source: jayavibhav/prompt-injection-safety
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Russian translated by Google Translate
score_ru_google - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-russian-gt.
