datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
russian-names
Russian Names with Popularity Scores
Description
This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.german-courts
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rusheeliyer/german-courts.ru_sentiment_dataset
Dataset with sentiment of Russian text
Contains aggregated dataset of Russian texts from 6 datasets.
Labels meaning
0: NEUTRAL
1: POSITIVE
2: NEGATIVE
Datasets
Sentiment Analysis in Russian
Sentiments (positive, negative or neutral) of news in russian language from Kaggle competition.
Russian Language Toxic Comments
Small dataset with labeled comments from 2ch.hk and pikabu.ru.
Dataset of car reviews for machine learning (sentiment analysis)
Glazkova A.… See the full description on the dataset page: https://huggingface.co/datasets/MonoHime/ru_sentiment_dataset.Srilanka-vegetable-prices
Sri Lanka Daily Price Dataset Pipeline
Automates extraction of selected items from the CBSL Daily Price Report PDF and appends them to a long-format CSV. It also enriches each date with rainfall for Nuwara Eliya and Polonnaruwa using Open-Meteo.
Output Schema
Columns in data/price_dataset.csv:
date (YYYY-MM-DD)
item
unit
retail_pettah
retail_dambulla
retail_narahenpita
wholesale_pettah
wholesale_dambulla
rainfall_nuwara_eliya_mm
rainfall_polonnaruwa_mm
source_pdf… See the full description on the dataset page: https://huggingface.co/datasets/Rusiru-erandaka/Srilanka-vegetable-prices.bioleaflets-biomedical-ner
Dataset Card for BioLeaflets Dataset
Dataset Summary
BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website.
Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately.
This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.russian-business-registries
Russian Business Registries — Aggregated Statistics
Aggregated, ready-to-analyse slices of Russian state registers. Every figure
comes from an official open-data source; nothing here is modelled, imputed or
estimated. Individual companies are not published — only aggregates, with one
deliberate exception described below.
Собрано из открытых данных российских госреестров. Все цифры — из официальных
источников, без моделирования и досчётов. Публикуются агрегаты, не сведения
об… See the full description on the dataset page: https://huggingface.co/datasets/Roman-Kpro/russian-business-registries.Russian_bank_reviews
Dataset Card for bank reviews dataset
Dataset Summary
The dataset is collected from the banki.ru website.
It contains customer reviews of various banks. In total, the dataset contains 12399 reviews.
The dataset is suitable for sentiment classification.
The dataset contains this fields - bank name, username, review title, review text, review time, number of views,
number of comments, review rating set by the user, as well as ratings for special categories… See the full description on the dataset page: https://huggingface.co/datasets/Romjiik/Russian_bank_reviews.russian_gec
📝 Dataset Card for Russian Grammar Error-Correction (25 362 sentence pairs)
A compact, high-quality corpus of Russian sentences with grammatical errors aligned to their human‑corrected counterparts. Ideal for training and benchmarking grammatical error‑correction (GEC) models, writing assistants, and translation post‑editing.
✨ Dataset Summary
Metric
Value
Sentence pairs
25 362
Avg. tokens / sentence
≈ 12
File size
~5 MB (CSV, UTF‑8)
Error types… See the full description on the dataset page: https://huggingface.co/datasets/dreuxx26/russian_gec.russian-speech-recognition-dataset
Russian Speech Dataset for recognition task
Dataset comprises 338 hours of telephone dialogues in Russian, collected from 460 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/russian-speech-recognition-dataset.russian-20th-century-bigrams
Русскоязычные биграммы XX века
Общие замечания
В этом датасете содержатся преобразованные в более подходящий для исследования вид биграммы на русском языке и их частотности из коллекции Google Ngrams с 1918 до 2010 года. Такой выбор обусловлен диапазоном дат, в рамках которого в русском языке соблюдается современный орфографический режим. Дореформенная орфография хуже распознается системами OCR, с ней невозможно работать as is современными средствами обработки… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-20th-century-bigrams.keyphrase_extraction_russianA dataset for extracting or generating keywords for scientific texts in Russian. It contains annotations of scientific articles from four domains: mathematics and computer science (a fragment of the Keyphrases CS&Math Russian dataset), history, linguistics, and medicine. For more details about the dataset please refer the original paper: https://arxiv.org/abs/2409.10640
Dataset Structure
abstract, an abstract in a string format;
keywords, a list of keyphrases provided by… See the full description on the dataset page: https://huggingface.co/datasets/aglazkova/keyphrase_extraction_russian.lezgi-rus-azer-corpus
Neural machine translation system for Lezgian, Russian and Azerbaijani languages
We release the first neural machine translation system for translation between Russian, Azerbaijani and the endangered Lezgian languages, as well as monolingual and parallel datasets collected and aligned for training and evaluating the system.
Parallel corpus sources:
Bible: parsed from https://bible.com, aligned by verse numbers (merged if necessary).
"Oriental translation" (CARS) version… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/lezgi-rus-azer-corpus.repa
Dataset Card for REPA
Image Source
Dataset Description
REPA is a Russian language dataset which consists of 1k user queries categorized into nine types, along with responses from six open-source instruction-finetuned Russian LLMs. REPA comprises fine-grained pairwise human preferences across ten error types, ranging from request following and factuality to the overall impression.
Each data instance consists of a query and two LLM responses manually annotated to… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/repa.russian-spell-correctionsRusFinQABenchmark
RusFinQABenchmark
Результаты оценки шести больших языковых моделей на датасете RuFinQA с использованием системы метрик FinCoT-Eval.
Модели
gemma2:9b
qwen2.5:7b
deepseek-r1:7b
phi3:3.8b
llama3.1:8b
aya:8b
Метрики
FinCoT-Eval: FAA, OT, CSPS, NEPS, FCS, Composite
Текстовые: COMET, BERTScore, ROUGE, BLEU
Структура файлов
evaluation.csv — построчные метрики для 6000 генераций (1000 вопросов × 6 моделей)
summary.csv — агрегированные… See the full description on the dataset page: https://huggingface.co/datasets/RusNLPWorld/RusFinQABenchmark.The_Secret_of_Third_Planet_Rus_Lezkomi-russian-parallel-corpora
Source Datasets
1 - news from the website of the Komi administration (https://rkomi.ru/)
2 - Komi media library (http://videocorpora.ru/)
3 - Millet porridge by Ivan Toropov (adaptation)
Authors
Shilova Nadezhda
Chernousov Georgy
rus-scifact-qrelsnova_sft
Nova SFT
This is a merged dataset from:
OpenHermes 2.5
AllenAI Tulu-3-SFT
Version 1, made on 2nd March 2025
russian_oil_gas_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP
Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP
Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.depflow_dataset
Depflow Sinhala and English Mix Coded Depressive Social Media Content Dataset
Ultracleaned-V1russian-speech-dataset
Russian Speech Dataset
The Russian Speech Dataset is a structured speech audio dataset designed to deliver high-quality audio data for machine learning and AI-driven voice systems. It includes 91 hours of audio data distributed across 641 files, provided in MP3 and WAV formats with a total size of 307 MB.
This well-organized audio dataset ensures balanced voice data, with 50% female and 50% male speakers, and a broad age distribution from 18 to 50+ years. The dataset language is… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/russian-speech-dataset.russian_llm_response_chatgpt_distill
LLM Usage in Russian (Distilled Dataset)
Dataset Summary
LLM Usage RU Dataset is a synthetic dataset of 50,000 Russian-language human–LLM interaction logs. Each sample includes a user query, the LLM's response, timestamp, user feedback, and session metadata. The dataset was generated to explore how large language models perform in Russian — a language that tends to receive less training coverage than English.
The queries and responses were distilled from GPT-4-turbo… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/russian_llm_response_chatgpt_distill.rus-nfcorpus-qrelsrussian-surnames
Russian Surnames
Description
This dataset contains over 300,000 Russian surnames with gender classification (m, f, u).
For a dataset of Russian given names, see Russian Names with Popularity Scores.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("rustemgareev/russian-surnames", split='train')
# Print the first example
print(dataset[0])
Example Output:
{… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-surnames.Lyrical_rus2eng_ORPOv5.1_SongsPoems_MeteredTranslations_csv
Meaning+Meter-Matched Russian & Soviet Poems + Songs
Manually Translated by a Poet-Translator from Russian to English
Translations herein faithfully adapt the Source Lyrics' Metered/Rhythmic/Rhyming Patterns
NEWLY EDITED VARIANT 5.1: 1775 rows/items
Re-balanced, refined, standardized, and substantially expanded.
CSV version
Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically reproduce source texts'… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_rus2eng_ORPOv5.1_SongsPoems_MeteredTranslations_csv.tyv-rus-200k
tyv-rus-200k data card
This data was collected via www.tyvan.ru platform by linguists, scientists, journalists, volunteers, etc.
Actually here 296k rows. Almost 300k
Dataset Details
Dataset Description
Curated by: Ali Kuzhuget (tech and data), Ondar Choygan (data) contributors
Language(s) (NLP): Tyvan (Tuvan), Russian
License:: CC BY 4.0.
Below is the brief information about the languages
Language
Language code on the website
ISO 639-3
Glottolog… See the full description on the dataset page: https://huggingface.co/datasets/Agisight/tyv-rus-200k.
