CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ruslanmv /sports-trends-dataset ⚽🏀🎾🏏 Sports-Trends Dataset A leakage-safe, multi-sport match data lake — raw fixtures → engineered features → training splits. The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day. TL;DR — A continuously-updated, medallion-architecture data lake for football, basketball, tennis and cricket: immutable raw ingests, cleaned/standardized layers, an engineered feature store, and ready-to-train chronological splits in… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/sports-trends-dataset.tabulartabular-classificationn<1K5 likes4.8k downloads6h agoHugging Face02ruslanmv /ai-medical-chatbot AI Medical Chatbot Dataset This is an experimental Dataset designed to run a Medical Chatbot It contains at least 250k dialogues between a Patient and a Doctor. Playground ChatBot ruslanmv/AI-Medical-Chatbot For furter information visit the project here: https://github.com/ruslanmv/ai-medical-chatbot text100K<n<1M253 likes4.8k downloads2y agoHugging Face03RussianNLP /wikiomnia Dataset Card for "Wikiomnia" Dataset Summary We present the WikiOmnia dataset, a new publicly available set of QA-pairs and corresponding Russian Wikipedia article summary sections, composed with a fully automated generative pipeline. The dataset includes every available article from Wikipedia for the Russian language. The WikiOmnia pipeline is available open-source and is also tested for creating SQuAD-formatted QA on other domains, like news texts, fiction, and social… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/wikiomnia.question-answering1M<n<10M18 likes3.7k downloads3y agoHugging Face04Rusa /musixaudion<1K0 likes3.1k downloads8mo agoHugging Face05r1v3r /multi_SWE_Bench_Rust multi_SWE_Bench_Rust 数据集描述... textn<1K1 likes2.6k downloads1y agoHugging Face06RussianNLP /coat Dataset Card for CoAT🧥 Dataset Description CoAT🧥 (Corpus of Artificial Texts) is a large-scale corpus for Russian, which consists of 246k human-written texts from publicly available resources and artificial texts generated by 13 neural models, varying in the number of parameters, architecture choices, pre-training objectives, and downstream applications. Each model is fine-tuned for one or more of six natural language generation tasks, ranging from paraphrase generation… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/coat.tabulartext-classification100K<n<1M4 likes1.7k downloads10mo agoHugging Face07irlspbru /RusLawOD The Russian Legislative Corpus, 1991–2026 Russian primary and secondary legislation corpus covering laws of Russian Federation, decrees by the President of RF, regulations by the government published as of July, 2026. The corpus collects all 308,056 texts (198,777,737 tokens) of non-secret federal regulations and acts, along with their metadata. The corpus has two versions: the original text with minimal preprocessing and a version prepared for linguistic analysis with… See the full description on the dataset page: https://huggingface.co/datasets/irlspbru/RusLawOD.text100K<n<1M18 likes1.2k downloads15d agoHugging Face08ammarnasr /the-stack-rust-clean Dataset 1: TheStack - Rust - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language. Target Language: Rust Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.tabulartext-generation100K<n<1M23 likes1.2k downloads2y agoHugging Face09ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1.1k downloads1y agoHugging Face10Den4ikAI /russian_dialoguesДатасет русских диалогов собранных с Telegram чатов. Диалоги имеют разметку по релевантности. Также были сгенерированы негативные примеры с помощью перемешивания похожих ответов. Количество диалогов - 2 миллиона Формат датасета: { 'question': 'Привет', 'answer': 'Привет, как дела?' 'relevance': 1 } Программа парсинга: https://github.com/Den4ikAI/telegram_chat_parser Citation: @MISC{russian_instructions, author = {Denis Petrov}, title = {Russian dialogues… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_dialogues.text1M<n<10M50 likes1.1k downloads4y agoHugging Face11esimijoq /Kazakh-Russian-Child-Directed-Speech-Corpus Kazakh-Russian Child-Directed Language Corpus Dataset Description The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization. The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts. This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.texttext-generation1M<n<10M0 likes1.1k downloads15d agoHugging Face12changliu8541 /assemblage-rust Assemblage-Rust Consider using your coding agent to test, download and process the data, repository size is 1.5T and it is highly likely you only want a portion of it, but do remember to check the agent outputs. Produced by Assemblage, a distributed binary-corpus generator, a cite would be greatly appreciated! Teh dataset contains127,165 compiled Rust binaries from 77,004 builds of 10,766 permissively licensed GitHub repositories, each paired with DWARF-derived function and… See the full description on the dataset page: https://huggingface.co/datasets/changliu8541/assemblage-rust.10K<n<100K0 likes1k downloads26d agoHugging Face13Slait /russia_voicesRussian voices for train AI. ♀ Male and ♂ Female Male voices - 497 pcs. Female voices - 244 pcs. Prepared for training fish-speech The author is not responsible for the votes. Use at your own risk. license: apache-2.0 task_categories: - zero-shot-classification language: - ru size_categories: - 1B<n<10B audio12 likes1k downloads1y agoHugging Face14Wholesomeisland /rust-the-stack-v2text1M<n<10M0 likes975 downloads5mo agoHugging Face15AdoCleanCode /SPEEED_s3_words_russian_0k-300ktext100K<n<1M0 likes917 downloads7mo agoHugging Face16istupakov /russian_librispeech Russian LibriSpeech (RuLS) Identifier: SLR96 from openslr.org Summary: This dataset is based on LibriVox audiobooks Category: Speech License: The dataset is Public Domain in the USA. About this resource: Russian LibriSpeech (RuLS) dataset is based on LibriVox's public domain audio books (see BOOKS.TXT for the list of included books) and contains about 98 hours of audio data. audioautomatic-speech-recognition10K<n<100K6 likes830 downloads1y agoHugging Face17RussianNLP /russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills - detection of natural language inference, commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating models and an overall leaderboard of transformer models for the Russian language.text-classification100K<n<1M35 likes779 downloads3y agoHugging Face18nevmenandr /russian-old-orthography-ocr Basic Description Dataset contains source images and human-readable extracted texts. All texts were published in Russia in the 19th century and written using pre-reform orthography. The dataset is designed to train and evaluate optical character recognition systems for texts published in Russian before the orthographic reform (1917). Data structure For each text there is a file with its image and the text corresponding to this image. The names of these files are the same… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-old-orthography-ocr.image100K<n<1M7 likes763 downloads2y agoHugging Face19AdoCleanCode /SPEEED_s3_words_russian_300k-600ktext100K<n<1M0 likes739 downloads7mo agoHugging Face20r1v3r /multiswe_rustbenchtextn<1K1 likes707 downloads1y agoHugging Face21RussianNLP /rublimp RuBLiMP Dataset Description RuBLiMP, or Russian Benchmark of Linguistic Minimal Pairs, is the first diverse and large-scale benchmark of minimal pairs in Russian. RuBLiMP includes 45k minimal pairs of sentences that differ in grammaticality and isolate morphological, syntactic, or semantic phenomena. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/rublimp.tabular10K<n<100K4 likes699 downloads1y agoHugging Face22Den4ikAI /russian_dialogues_2 Den4ikAI/russian_dialogues_2 Датасет русских диалогов для обучения диалоговых моделей. Количество диалогов - 1.6 миллиона Формат датасета: { 'sample': ['Привет', 'Привет', 'Как дела?'] } Citation: @MISC{russian_instructions, author = {Denis Petrov}, title = {Russian context dialogues dataset for conversational agents}, url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2}, year = 2023 } texttext-generation1M<n<10M18 likes641 downloads2y agoHugging Face23RussianNLP /tapeThe Winograd schema challenge composes tasks with syntactic ambiguity, which can be resolved with logic and reasoning (Levesque et al., 2012). The texts for the Winograd schema problem are obtained using a semi-automatic pipeline. First, lists of 11 typical grammatical structures with syntactic homonymy (mainly case) are compiled. For example, two noun phrases with a complex subordinate: 'A trinket from Pompeii that has survived the centuries'. Requests corresponding to these constructions are submitted in search of the Russian National Corpus, or rather its sub-corpus with removed homonymy. In the resulting 2+k examples, homonymy is removed automatically with manual validation afterward. Each original sentence is split into multiple examples in the binary classification format, indicating whether the homonymy is resolved correctly or not.text-classification1K<n<10K10 likes585 downloads2y agoHugging Face24AlienKevin /Multi-SWE-smith-Rust-GLM-4.6-trajectoriestextn<1K0 likes578 downloads9mo agoHugging Face25gaianet /learn-rust Knowledge base from the Rust books Gaia node setup instructions See the Gaia node getting started guide gaianet init --config https://raw.githubusercontent.com/GaiaNet-AI/node-configs/main/llama-3.1-8b-instruct_rustlang/config.json gaianet start with the full 128k context length of Llama 3.1 gaianet init --config https://raw.githubusercontent.com/GaiaNet-AI/node-configs/main/llama-3.1-8b-instruct_rustlang/config.fullcontext.json gaianet start Steps to create… See the full description on the dataset page: https://huggingface.co/datasets/gaianet/learn-rust.1K<n<10K11 likes559 downloads2y agoHugging Face26Den4ikAI /russian_instructions_2June 10: Почищены криво переведенные примеры кода Добавлено >50000 человеческих примеров QA и инструкций Обновленная версия русского датасета инструкций и QA. Улучшения: 1. Увеличен размер с 40 мегабайт до 130 (60к сэмплов - 200к) 2. Улучшено качество перевода. Структура датасета: { "sample":[ "Как я могу улучшить свою связь между телом и разумом?", "Начните с разработки регулярной практики осознанности. 2. Обязательно практикуйте баланс на нескольких уровнях: физическом… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_instructions_2.text100K<n<1M27 likes555 downloads3y agoHugging Face27HumynLabs /Russian_Documents_Dataset_PDF Russian Documents Dataset (PDF) This dataset contains a curated collection of Russian-language documents in PDF format. The corpus includes books, academic papers, government publications, articles, and educational materials written in Russian. It is designed to support AI research in OCR, document understanding, and multilingual text recognition. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Russian_Documents_Dataset_PDF.documentn<1K1 likes532 downloads11mo agoHugging Face28kijjjj /audio_data_russian_backup Dataset Audio Russian Backup This is a backup dataset with Russian audio data, split into train_0 to train_49 for tasks like text-to-speech, speech recognition, and speaker identification. Features text: Audio transcription (string). speaker_name: Speaker identifier (string). audio: Audio file. Usage Load the dataset like this: from datasets import load_dataset dataset = load_dataset("kijjjj/audio_data_russian_backup", split="train_0") # Or any train_X… See the full description on the dataset page: https://huggingface.co/datasets/kijjjj/audio_data_russian_backup.audiotext-to-speech100K<n<1M0 likes506 downloads1y agoHugging Face29P90-RushB /AgentArk AgentArk Assets Versioned Unity runtimes, task Mods, and multimodal replay/evaluation records. Environment compatibility / 环境兼容性 A task's labeled env version is its minimum supported environment version. That version and all later versions are compatible, with no upper version bound. 任务标注的 env 版本是最低环境版本;该版本及之后的所有版本均兼容,没有版本上限。 Task artifact label / 任务标注 Compatible environments / 可用环境 env-1.0.3 >=1.0.3: 1.0.3, 1.0.4, 1.0.5, and all later versions… See the full description on the dataset page: https://huggingface.co/datasets/P90-RushB/AgentArk.reinforcement-learning0 likes504 downloads2d agoHugging Face30r1v3r /rustbenchtextn<1K0 likes486 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.