CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rustemgareev /russian-names Russian Names with Popularity Scores Description This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025. Usage The dataset can be loaded using the Hugging Face datasets library. from datasets import load_dataset dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.tabularother10K<n<100K0 likes428 downloads1y agoHugging Face02rusheeliyer /german-courts Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rusheeliyer/german-courts.text1K<n<10K1 likes339 downloads3y agoHugging Face03MonoHime /ru_sentiment_dataset Dataset with sentiment of Russian text Contains aggregated dataset of Russian texts from 6 datasets. Labels meaning 0: NEUTRAL 1: POSITIVE 2: NEGATIVE Datasets Sentiment Analysis in Russian Sentiments (positive, negative or neutral) of news in russian language from Kaggle competition. Russian Language Toxic Comments Small dataset with labeled comments from 2ch.hk and pikabu.ru. Dataset of car reviews for machine learning (sentiment analysis) Glazkova A.… See the full description on the dataset page: https://huggingface.co/datasets/MonoHime/ru_sentiment_dataset.tabular100K<n<1M13 likes329 downloads5y agoHugging Face04Rusiru-erandaka /Srilanka-vegetable-prices Sri Lanka Daily Price Dataset Pipeline Automates extraction of selected items from the CBSL Daily Price Report PDF and appends them to a long-format CSV. It also enriches each date with rainfall for Nuwara Eliya and Polonnaruwa using Open-Meteo. Output Schema Columns in data/price_dataset.csv: date (YYYY-MM-DD) item unit retail_pettah retail_dambulla retail_narahenpita wholesale_pettah wholesale_dambulla rainfall_nuwara_eliya_mm rainfall_polonnaruwa_mm source_pdf… See the full description on the dataset page: https://huggingface.co/datasets/Rusiru-erandaka/Srilanka-vegetable-prices.tabular1K<n<10K0 likes146 downloads15d agoHugging Face05ruslan /bioleaflets-biomedical-ner Dataset Card for BioLeaflets Dataset Dataset Summary BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website. Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately. This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.texttext-generation1K<n<10K4 likes112 downloads4y agoHugging Face06Roman-Kpro /russian-business-registries Russian Business Registries — Aggregated Statistics Aggregated, ready-to-analyse slices of Russian state registers. Every figure comes from an official open-data source; nothing here is modelled, imputed or estimated. Individual companies are not published — only aggregates, with one deliberate exception described below. Собрано из открытых данных российских госреестров. Все цифры — из официальных источников, без моделирования и досчётов. Публикуются агрегаты, не сведения об… See the full description on the dataset page: https://huggingface.co/datasets/Roman-Kpro/russian-business-registries.tabulartabular-classification10K<n<100K0 likes109 downloads9d agoHugging Face07Romjiik /Russian_bank_reviews Dataset Card for bank reviews dataset Dataset Summary The dataset is collected from the banki.ru website. It contains customer reviews of various banks. In total, the dataset contains 12399 reviews. The dataset is suitable for sentiment classification. The dataset contains this fields - bank name, username, review title, review text, review time, number of views, number of comments, review rating set by the user, as well as ratings for special categories… See the full description on the dataset page: https://huggingface.co/datasets/Romjiik/Russian_bank_reviews.tabulartext-classification10K<n<100K7 likes106 downloads3y agoHugging Face08dreuxx26 /russian_gec 📝 Dataset Card for Russian Grammar Error-Correction (25 362 sentence pairs) A compact, high-quality corpus of Russian sentences with grammatical errors aligned to their human‑corrected counterparts. Ideal for training and benchmarking grammatical error‑correction (GEC) models, writing assistants, and translation post‑editing. ✨ Dataset Summary Metric Value Sentence pairs 25 362 Avg. tokens / sentence ≈ 12 File size ~5 MB (CSV, UTF‑8) Error types… See the full description on the dataset page: https://huggingface.co/datasets/dreuxx26/russian_gec.text10K<n<100K3 likes105 downloads1y agoHugging Face09UniDataPro /russian-speech-recognition-dataset Russian Speech Dataset for recognition task Dataset comprises 338 hours of telephone dialogues in Russian, collected from 460 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/russian-speech-recognition-dataset.textn<1K7 likes99 downloads1mo agoHugging Face10nevmenandr /russian-20th-century-bigrams Русскоязычные биграммы XX века Общие замечания В этом датасете содержатся преобразованные в более подходящий для исследования вид биграммы на русском языке и их частотности из коллекции Google Ngrams с 1918 до 2010 года. Такой выбор обусловлен диапазоном дат, в рамках которого в русском языке соблюдается современный орфографический режим. Дореформенная орфография хуже распознается системами OCR, с ней невозможно работать as is современными средствами обработки… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-20th-century-bigrams.tabular10M<n<100M1 likes92 downloads11mo agoHugging Face11aglazkova /keyphrase_extraction_russianA dataset for extracting or generating keywords for scientific texts in Russian. It contains annotations of scientific articles from four domains: mathematics and computer science (a fragment of the Keyphrases CS&Math Russian dataset), history, linguistics, and medicine. For more details about the dataset please refer the original paper: https://arxiv.org/abs/2409.10640 Dataset Structure abstract, an abstract in a string format; keywords, a list of keyphrases provided by… See the full description on the dataset page: https://huggingface.co/datasets/aglazkova/keyphrase_extraction_russian.texttext-generation10K<n<100K5 likes84 downloads11mo agoHugging Face12AlidarAsvarov /lezgi-rus-azer-corpus Neural machine translation system for Lezgian, Russian and Azerbaijani languages We release the first neural machine translation system for translation between Russian, Azerbaijani and the endangered Lezgian languages, as well as monolingual and parallel datasets collected and aligned for training and evaluating the system. Parallel corpus sources: Bible: parsed from https://bible.com, aligned by verse numbers (merged if necessary). "Oriental translation" (CARS) version… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/lezgi-rus-azer-corpus.text100K<n<1M5 likes61 downloads2y agoHugging Face13RussianNLP /repa Dataset Card for REPA Image Source Dataset Description REPA is a Russian language dataset which consists of 1k user queries categorized into nine types, along with responses from six open-source instruction-finetuned Russian LLMs. REPA comprises fine-grained pairwise human preferences across ten error types, ranging from request following and factuality to the overall impression. Each data instance consists of a query and two LLM responses manually annotated to… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/repa.text1K<n<10K4 likes60 downloads2y agoHugging Face14chainiy /russian-spell-correctionstext1K<n<10K1 likes58 downloads2d agoHugging Face15RusNLPWorld /RusFinQABenchmark RusFinQABenchmark Результаты оценки шести больших языковых моделей на датасете RuFinQA с использованием системы метрик FinCoT-Eval. Модели gemma2:9b qwen2.5:7b deepseek-r1:7b phi3:3.8b llama3.1:8b aya:8b Метрики FinCoT-Eval: FAA, OT, CSPS, NEPS, FCS, Composite Текстовые: COMET, BERTScore, ROUGE, BLEU Структура файлов evaluation.csv — построчные метрики для 6000 генераций (1000 вопросов × 6 моделей) summary.csv — агрегированные… See the full description on the dataset page: https://huggingface.co/datasets/RusNLPWorld/RusFinQABenchmark.tabularquestion-answering1K<n<10K0 likes56 downloads28d agoHugging Face16leks-forever /The_Secret_of_Third_Planet_Rus_Leztexttranslationn<1K1 likes53 downloads2y agoHugging Face17Horeknad /komi-russian-parallel-corpora Source Datasets 1 - news from the website of the Komi administration (https://rkomi.ru/) 2 - Komi media library (http://videocorpora.ru/) 3 - Millet porridge by Ivan Toropov (adaptation) Authors Shilova Nadezhda Chernousov Georgy texttranslation10K<n<100K2 likes51 downloads3y agoHugging Face18kaengreg /rus-scifact-qrelstabular1K<n<10K0 likes49 downloads2y agoHugging Face19Rushfoxsea /nova_sft Nova SFT This is a merged dataset from: OpenHermes 2.5 AllenAI Tulu-3-SFT Version 1, made on 2nd March 2025 tabular100K<n<1M1 likes48 downloads2y agoHugging Face20ScoutieAutoML /russian_oil_gas_news_telegram_dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.tabulartext-classification10K<n<100K7 likes46 downloads2y agoHugging Face21ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes46 downloads2y agoHugging Face22fahim-ling /Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLPgated Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.texttranslationn<1K1 likes46 downloads22d agoHugging Face23rusiruwijethilake /depflow_dataset Depflow Sinhala and English Mix Coded Depressive Social Media Content Dataset text1K<n<10K1 likes44 downloads3y agoHugging Face24Rushfoxsea /Ultracleaned-V1tabular100K<n<1M1 likes43 downloads1y agoHugging Face25Speech-data /russian-speech-dataset Russian Speech Dataset The Russian Speech Dataset is a structured speech audio dataset designed to deliver high-quality audio data for machine learning and AI-driven voice systems. It includes 91 hours of audio data distributed across 641 files, provided in MP3 and WAV formats with a total size of 307 MB. This well-organized audio dataset ensures balanced voice data, with 50% female and 50% male speakers, and a broad age distribution from 18 to 50+ years. The dataset language is… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/russian-speech-dataset.audioautomatic-speech-recognitionn<1K0 likes42 downloads6mo agoHugging Face26ZennyKenny /russian_llm_response_chatgpt_distill LLM Usage in Russian (Distilled Dataset) Dataset Summary LLM Usage RU Dataset is a synthetic dataset of 50,000 Russian-language human–LLM interaction logs. Each sample includes a user query, the LLM's response, timestamp, user feedback, and session metadata. The dataset was generated to explore how large language models perform in Russian — a language that tends to receive less training coverage than English. The queries and responses were distilled from GPT-4-turbo… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/russian_llm_response_chatgpt_distill.text10K<n<100K2 likes41 downloads1y agoHugging Face27kaengreg /rus-nfcorpus-qrelstext100K<n<1M0 likes39 downloads2y agoHugging Face28rustemgareev /russian-surnames Russian Surnames Description This dataset contains over 300,000 Russian surnames with gender classification (m, f, u). For a dataset of Russian given names, see Russian Names with Popularity Scores. Usage The dataset can be loaded using the Hugging Face datasets library. from datasets import load_dataset # Load the dataset dataset = load_dataset("rustemgareev/russian-surnames", split='train') # Print the first example print(dataset[0]) Example Output: {… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-surnames.text100K<n<1M0 likes39 downloads1y agoHugging Face29AlekseyCalvin /Lyrical_rus2eng_ORPOv5.1_SongsPoems_MeteredTranslations_csv Meaning+Meter-Matched Russian & Soviet Poems + Songs Manually Translated by a Poet-Translator from Russian to English Translations herein faithfully adapt the Source Lyrics' Metered/Rhythmic/Rhyming Patterns NEWLY EDITED VARIANT 5.1: 1775 rows/items Re-balanced, refined, standardized, and substantially expanded. CSV version Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically reproduce source texts'… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_rus2eng_ORPOv5.1_SongsPoems_MeteredTranslations_csv.texttranslation1K<n<10K0 likes39 downloads11mo agoHugging Face30Agisight /tyv-rus-200k tyv-rus-200k data card This data was collected via www.tyvan.ru platform by linguists, scientists, journalists, volunteers, etc. Actually here 296k rows. Almost 300k Dataset Details Dataset Description Curated by: Ali Kuzhuget (tech and data), Ondar Choygan (data) contributors Language(s) (NLP): Tyvan (Tuvan), Russian License:: CC BY 4.0. Below is the brief information about the languages Language Language code on the website ISO 639-3 Glottolog… See the full description on the dataset page: https://huggingface.co/datasets/Agisight/tyv-rus-200k.texttranslation100K<n<1M0 likes37 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.