CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RussianNLP /wikiomnia Dataset Card for "Wikiomnia" Dataset Summary We present the WikiOmnia dataset, a new publicly available set of QA-pairs and corresponding Russian Wikipedia article summary sections, composed with a fully automated generative pipeline. The dataset includes every available article from Wikipedia for the Russian language. The WikiOmnia pipeline is available open-source and is also tested for creating SQuAD-formatted QA on other domains, like news texts, fiction, and social… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/wikiomnia.question-answering1M<n<10M18 likes3.9k downloads3y agoHugging Face02RussianNLP /russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills - detection of natural language inference, commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating models and an overall leaderboard of transformer models for the Russian language.text-classification100K<n<1M35 likes797 downloads3y agoHugging Face03RussianNLP /tapeThe Winograd schema challenge composes tasks with syntactic ambiguity, which can be resolved with logic and reasoning (Levesque et al., 2012). The texts for the Winograd schema problem are obtained using a semi-automatic pipeline. First, lists of 11 typical grammatical structures with syntactic homonymy (mainly case) are compiled. For example, two noun phrases with a complex subordinate: 'A trinket from Pompeii that has survived the centuries'. Requests corresponding to these constructions are submitted in search of the Russian National Corpus, or rather its sub-corpus with removed homonymy. In the resulting 2+k examples, homonymy is removed automatically with manual validation afterward. Each original sentence is split into multiple examples in the binary classification format, indicating whether the homonymy is resolved correctly or not.text-classification1K<n<10K10 likes452 downloads2y agoHugging Face04wwewtech /russian-it-community-corpus 📦 Russian IT Community Corpus (RICC) Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture. The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.tabulartext-generation1M<n<10M1 likes331 downloads16d agoHugging Face05drkolesnikov /russian-nmo-medical-mcq Russian NMO Medical MCQ Choose language / Выберите язык: Русский | English Русский Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа. В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA. Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.tabularquestion-answering1M<n<10M2 likes174 downloads4mo agoHugging Face06attn-signs /russian-easy-instructions Easy Russian Instructions Fast instructions in conversational format, generaly used for injecting general knowledge.Dataset contains comprehensive instructions for easy tasks and general question answering Contents: Wikipedia / Factological knowledge History knowledge Basic programming understanding Basic math understanding Basic physics understanding Basic geography knowledge Basic biology knowledge Format: Formatted for instruction-tuning / instruction… See the full description on the dataset page: https://huggingface.co/datasets/attn-signs/russian-easy-instructions.textquestion-answering10K<n<100K1 likes111 downloads2y agoHugging Face07issai /ARC_Kazakh_Russian Dataset Summary These are the machine-translated Kazakh and Russian versions of the ARC (AI2 Reasoning Challenge) dataset (test set). ARC is specifically designed to focus on genuine reasoning abilities, containing elementary and middle-school science questions that require more than simple retrieval or word-matching to solve. These Kazakh and Russian versions serve as a benchmark for evaluating how well models can perform complex reasoning and apply scientific knowledge when… See the full description on the dataset page: https://huggingface.co/datasets/issai/ARC_Kazakh_Russian.textquestion-answering1K<n<10K0 likes91 downloads1mo agoHugging Face08issai /GPQA_Kazakh_Russian Dataset Summary These are the machine-translated Kazakh and Russian versions of the GPQA (Graduate-Level Google-Proof Q&A Benchmark) dataset. These datasets are used to test the world knowledge and problem-solving capabilities of large language models across a vast range of subjects in the Kazakh and Russian languages. Unlike general knowledge benchmarks, GPQA consists of extremely challenging science questions (biology, physics, and chemistry) written by experts. These… See the full description on the dataset page: https://huggingface.co/datasets/issai/GPQA_Kazakh_Russian.textquestion-answering1K<n<10K0 likes73 downloads1mo agoHugging Face09issai /MMLU-Pro_Kazakh_Russian Dataset Summary These are the machine-translated Kazakh and Russian versions of the MMLU-Pro (Massive Multitask Language Understanding Pro) dataset (test set). These datasets are used to test the world knowledge and problem-solving capabilities of large language models across a vast range of subjects in the Kazakh and Russian languages. As an enhanced version of the original MMLU, it serves as a more rigorous benchmark for evaluating how well models understand complex academic… See the full description on the dataset page: https://huggingface.co/datasets/issai/MMLU-Pro_Kazakh_Russian.tabularquestion-answering10K<n<100K2 likes58 downloads1mo agoHugging Face10MLNavigator /russian-retrievalBased on Sberquad Answer converted to human affordable answer. Context augmented with some pices of texts from wiki accordant to text on tematic and keywords. This dataset cold be used for training retrieval LLM models or modificators for ability of LLM to retrieve target information from collection of tematic related texts. Dataset has version with SOURCE data for generating answer with specifing source document for right answer. See file retrieval_dataset_src.jsonl Dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/MLNavigator/russian-retrieval.textquestion-answering10K<n<100K5 likes47 downloads2y agoHugging Face11issai /MMstar_Kazakh_Russian Dataset Summary These are the machine-translated Kazakh and Russian versions of the MMStar dataset. MMStar is an elite vision-language benchmark consisting of high-quality, non-redundant multimodal samples designed to evaluate 6 core capabilities and 18 detailed dimensions of Large Multimodal Models (LMMs). Kazakh and Russian versions serve as a benchmark for evaluating how well models perform advanced visual recognition and multi-step reasoning while avoiding "data leakage"… See the full description on the dataset page: https://huggingface.co/datasets/issai/MMstar_Kazakh_Russian.imagequestion-answering1K<n<10K1 likes41 downloads1mo agoHugging Face12issai /RealWorldQA_Kazakh_Russian Dataset Summary These are the machine-translated Kazakh and Russian versions of the RealWorldQA dataset. As a multimodal benchmark, RealWorldQA is specifically designed to evaluate the real-world spatial understanding and visual reasoning of models. Kazakh and Russian versions serve as a benchmark for evaluating how well models can understand physical environments, spatial relationships, and object attributes based on real-world images when prompted in Kazakh or Russian.… See the full description on the dataset page: https://huggingface.co/datasets/issai/RealWorldQA_Kazakh_Russian.imagequestion-answering1K<n<10K0 likes35 downloads1mo agoHugging Face13askatasuna /GSM8KInstruct_russiantextquestion-answering1K<n<10K0 likes31 downloads6mo agoHugging Face14omni-devel /RussianUltrachat100kЧасть датасета stingning/ultrachat, переведенная на русский язык. Заказать перевод вашего датасета на любой язык мира: https://t.me/I_Am_PyWebSol texttext-generation100K<n<1M4 likes27 downloads2y agoHugging Face15Expotion /russian-facts-qa RU Wikipedia QA Facts This dataset is based on articles from the Russian Wikipedia (CC BY-SA 4.0).The source articles were split into text chunks, then Gemma 3 4B was used to generate initial question–answer (QA) pairs, and Gemma 3 12B validated and refined them. Data Format Each record is stored in JSONL format (.jsonl), one object per line: {"q": "Какие страны подписали мирные договоры на Парижской конференции в 1947 году?", "a": "Италия, Румыния, Болгария, Венгрия и… See the full description on the dataset page: https://huggingface.co/datasets/Expotion/russian-facts-qa.textquestion-answering10K<n<100K1 likes13 downloads1y agoHugging Face16Katya-Iukhn /optic_QA_pdf_russianimagequestion-answering1K<n<10K1 likes11 downloads2y agoHugging Face17iamivan11 /russian-spell-correctiontextquestion-answering1K<n<10K0 likes4 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.