CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ammarnasr /the-stack-rust-clean Dataset 1: TheStack - Rust - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language. Target Language: Rust Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.tabulartext-generation100K<n<1M24 likes1.2k downloads2y agoHugging Face02ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1k downloads1y agoHugging Face03wwewtech /russian-it-community-corpus 📦 Russian IT Community Corpus (RICC) Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture. The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.tabulartext-generation1M<n<10M1 likes339 downloads16d agoHugging Face04adeshkin /khakas-russian-dict Khakas-Russian Dictionary (Dataset) Dataset Developer: Vasily AdeshkinContact for inquiries: adeshkin.vi@phystech.edu 📌 Important Notice & Citation When using this dataset, please be sure to cite this repository and the original dictionary: https://khakas.altaica.ru/dictionary/. Please note: Some optical character recognition (OCR) errors may still be present in the data. 🛠 Contribution & Authorship I am not the author of the original dictionary, but I have… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-dict.tabulartranslation10K<n<100K0 likes213 downloads5mo agoHugging Face05drkolesnikov /russian-nmo-medical-mcq Russian NMO Medical MCQ Choose language / Выберите язык: Русский | English Русский Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа. В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA. Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.tabularquestion-answering1M<n<10M2 likes183 downloads4mo agoHugging Face06NickIBrody /rust-code-suite NickIBrody/rust-code-suite Rust Code Suite is a public raw Rust source corpus built from open-source repositories and selected historical git revisions. Splits train.jsonl validation.jsonl test.jsonl Schema { "id": "owner/repo:path:chunk", "text": "...", "arch": "rust", "syntax": "rust", "kind": "rust-source", "repo": "owner/repo", "path": "src/lib.rs", "license": "GPL-2.0", "commit": "abcdef123456", "source_url":… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/rust-code-suite.tabulartext-generation1M<n<10M1 likes122 downloads4mo agoHugging Face07samedad /mem-and-russian-jokes-dataset 2 июля 2025 Добавлено новых уникальных анекдотов: 1311963 Количество записей в датасете: 521904 Добавил датасет анекдотов от IgorVolochay/russian_jokes не понимаю как я прошел мимо него, там очень много шуток, сам Евгений Ваганович очень часто обращался к этому датасету... Но вот и мое время пришло. Плюс обновил старый датасет и разбавил формулировки новыми начальными вопросами human_prompts = [ "Расскажи шутку", "Расскажи анекдот", "Знаешь какой-нибудь прикол?", "Скажи что-нибудь смешное"… See the full description on the dataset page: https://huggingface.co/datasets/samedad/mem-and-russian-jokes-dataset.tabulartext-generation100K<n<1M4 likes90 downloads1y agoHugging Face08jebkoralav /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/jebkoralav/ru-big-russian-dataset.tabulartext-generation1M<n<10M0 likes62 downloads10mo agoHugging Face09RusNLPWorld /RusFinQABenchmark RusFinQABenchmark Результаты оценки шести больших языковых моделей на датасете RuFinQA с использованием системы метрик FinCoT-Eval. Модели gemma2:9b qwen2.5:7b deepseek-r1:7b phi3:3.8b llama3.1:8b aya:8b Метрики FinCoT-Eval: FAA, OT, CSPS, NEPS, FCS, Composite Текстовые: COMET, BERTScore, ROUGE, BLEU Структура файлов evaluation.csv — построчные метрики для 6000 генераций (1000 вопросов × 6 моделей) summary.csv — агрегированные… See the full description on the dataset page: https://huggingface.co/datasets/RusNLPWorld/RusFinQABenchmark.tabularquestion-answering1K<n<10K0 likes56 downloads1mo agoHugging Face10ScoutieAutoML /russian_oil_gas_news_telegram_dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.tabulartext-classification10K<n<100K7 likes53 downloads2y agoHugging Face11ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes46 downloads2y agoHugging Face12adityabhushannagar /code-alchemy-rust CodeAlchemy Rust Rust-only derivative of open-alchemy/code-alchemy. It preserves the five training configs, two evaluation configs, original splits, row order, columns, values, and task/evaluation fields. Rows were selected from the source-native language labels: Rust and rust in training data and dev-eval rs in trace-eval Labels remain unchanged in the output. code-trace.external_packages is normalized to list<string> because source Parquet shards physically alternate between… See the full description on the dataset page: https://huggingface.co/datasets/adityabhushannagar/code-alchemy-rust.tabulartext-generation1M<n<10M0 likes45 downloads2mo agoHugging Face13TatarNLPWorld /tatar-english-russian-corpusgated Dataset Card: Tatar-English-Russian Parallel Corpus Dataset Details Dataset Description This dataset is a parallel corpus containing 14,983 sentences in three languages: Tatar, English, and Russian. It combines two distinct sources: KickItLikeShika/english-tatar-translation (7,746 entries) – an existing English-Tatar dataset with Russian translations added yasalma/tt-en-language-corpus (7,615 entries) – another English-Tatar corpus with newly added… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-english-russian-corpus.tabulartranslation10K<n<100K0 likes32 downloads1mo agoHugging Face14paiml /rust-cli-docs-corpus Rust CLI Documentation Corpus A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools. Dataset Description This corpus follows the Toyota Way principles and Popperian falsification methodology. Statistics Total entries: 80 Source repositories: 0 Validation score: 96/100 Supported Tasks Documentation Generation: Generate Rust doc comments from code signatures Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.tabulartext-generationn<1K0 likes30 downloads8mo agoHugging Face15BashkirNLPWorld /bashkir-russian-parallelgated Dataset Card for Bashkir-Russian Parallel Corpus Dataset Details Dataset Description Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart. The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.tabulartext-generation1M<n<10M0 likes27 downloads7d agoHugging Face16ScoutieAutoML /russian_jokes_with_vectors Description in English: The dataset is collected from Russian-language Telegram channels with jokes and anecdotes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_jokes_with_vectors.tabulartext-classification10K<n<100K5 likes25 downloads2y agoHugging Face17AtesiT /ru-stem-dialogues Russian STEM Educational Dialogues Описание Синтетический датасет русскоязычных учебных диалогов по STEM-темам (математика, физика, химия, биология, информатика, программирование, инженерия). Каждый диалог — реалистичное взаимодействие между пользователем (школьник / студент / профессионал) и ассистентом. Методология Модель: Qwen/Qwen2.5-7B-Instruct (4-bit NF4 quantization, bitsandbytes) Формат генерации: текстовый формат с разделителями… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-stem-dialogues.tabulartext-generationn<1K0 likes24 downloads2mo agoHugging Face18RusNLPWorld /RuFinQAgated 🏦 RusFinChain RusFinChain is a Russian benchmark for evaluating Large Language Models (LLMs) on financial analysis tasks with Ground-Truth Chain-of-Thought. 📊 Overview Total questions: 44,627 Task types: 7 Skills: 12 Difficulty levels: 3 Version: 3.3.0 Language: Russian 📖 Source Material & Data Licensing Foundational SourceThe methodological framework, financial formulas, problem typology, and a significant portion of the practical… See the full description on the dataset page: https://huggingface.co/datasets/RusNLPWorld/RuFinQA.tabularquestion-answering10K<n<100K0 likes21 downloads1mo agoHugging Face19ScoutieAutoML /russian_events_vectors Description in English: The dataset is collected from Russian-language Telegram channels with information about various events and happenings in Russian regions,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_events_vectors.tabulartext-classification10K<n<100K0 likes20 downloads2y agoHugging Face20Adilbai /kz-rus-articles-comprehensive 🇰🇿🇷🇺 Kazakh-Russian Articles Comprehensive Dataset A high-quality bilingual corpus for cross-lingual NLP research 📋 Dataset Overview The Kazakh-Russian Articles Comprehensive Dataset is a meticulously curated bilingual corpus designed to advance natural language processing research for Kazakh and Russian languages. This dataset addresses the critical need for high-quality parallel and comparable text resources in Central Asian language pairs, particularly… See the full description on the dataset page: https://huggingface.co/datasets/Adilbai/kz-rus-articles-comprehensive.tabulartranslationn<1K1 likes18 downloads1y agoHugging Face21ScoutieAutoML /scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification1K<n<10K0 likes14 downloads2y agoHugging Face22ScoutieAutoML /scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification10K<n<100K0 likes10 downloads2y agoHugging Face23RusNLPWorld /RuFinQA-litegated RuFinQA — A Massive Multi-Task Reasoning Benchmark for Russian Financial Report Understanding RuFinQA is a large-scale multi-task benchmark designed to evaluate the ability of language models to understand and reason over Russian statutory financial reports (Balance Sheet, Income Statement, Cash Flow Statement). It contains 36,330 question–answer pairs across 5 task types, automatically derived from real-world corporate accounting statements obtained from open government data… See the full description on the dataset page: https://huggingface.co/datasets/RusNLPWorld/RuFinQA-lite.tabularquestion-answering10K<n<100K0 likes9 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.