CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B39 likes46k downloads5d agoHugging Face02rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K274 likes25k downloads3y agoHugging Face03d0rj /LLaVA-OneVision-Data-ru LLaVA-OneVision-Data-ru Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate. Almost all datasets have been translated, except for the following: ["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"] Usage import datasets data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.imagetext-generation1M<n<10M4 likes5k downloads2y agoHugging Face04RUC-AIBOX /Llama-3-SynE-Dataset 📄 Report   |   💻 GitHub Repo 🔍 English  |  简体中文 Here is the continual pre-training dataset. The Llama-3-SynE model is available here. News 🌟🌟 2024/12/17: We released the code used for continual pre-training and data preparation. The code contains detailed documentation comments. ✨✨ 2024/08/12: We released the continual pre-training dataset. ✨✨ 2024/08/10: We released the Llama-3-SynE model. ✨ 2024/07/26: We released the technical report, welcome to check it… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Llama-3-SynE-Dataset.texttext-generation100M<n<1B10 likes3k downloads1y agoHugging Face05Rubin-Wei /MemoryDecoder-at-Scale-domain-data MemoryDecoder at Scale Domain Data This repository contains the domain-specific continued-pretraining (CPT) data, the tokenized and preprocessed datasets, and the aligned KNN distributions used by MemoryDecoder at Scale. Links Project Page: Memory Decoder at Scale GitHub Repository: LUMIA-Group/MemoryDecoder-at-Scale Paper: Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory The preprocessed datasets and KNN distributions in this repository use… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/MemoryDecoder-at-Scale-domain-data.text-generation1 likes1.8k downloads2mo agoHugging Face06sojuL /RubricHub_v1 RubricHub RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of… See the full description on the dataset page: https://huggingface.co/datasets/sojuL/RubricHub_v1.texttext-generation100K<n<1M270 likes1.6k downloads8mo agoHugging Face07IlyaGusev /ru_turbo_alpaca RuTurboAlpaca Dataset of ChatGPT-generated instructions in Russian. Code: rulm/self_instruct Code is based on Stanford Alpaca and self-instruct. 29822 examples Preliminary evaluation by an expert based on 400 samples: 83% of samples contain correct instructions 63% of samples have correct instructions and outputs Crowdsouring-based evaluation on 3500 samples: 90% of samples contain correct instructions 68% of samples have correct instructions and outputs Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_alpaca.text-generation10K<n<100K69 likes1.3k downloads3y agoHugging Face08d0rj /ru-fandom-wiki d0rj/ru-fandom-wiki Description A set of texts collected from the most popular Russian-language fandoms (65 fandoms) on fandom.com. The dump given on 25.10.2024-27.10.2024 collected using trafilatura library. All texts are in markdown format. License The license supports the license text on the source site - Creative Commons Attribution-ShareAlike 3.0 (Unported) (CC-BY-SA). texttext-classification100K<n<1M5 likes1.2k downloads2y agoHugging Face09ammarnasr /the-stack-rust-clean Dataset 1: TheStack - Rust - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language. Target Language: Rust Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.tabulartext-generation100K<n<1M24 likes1.2k downloads2y agoHugging Face10ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1.1k downloads1y agoHugging Face11esimijoq /Kazakh-Russian-Child-Directed-Speech-Corpus Kazakh-Russian Child-Directed Language Corpus Dataset Description The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization. The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts. This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.texttext-generation1M<n<10M0 likes1.1k downloads15d agoHugging Face12Arushhh /Llama-HybridDiffusion-processed-data-run1 Llama-HybridDiffusion processed training mixture — run 1 Built with Llama. This repository preserves the exact Hugging Face Dataset.save_to_disk Arrow snapshot used by run 1 of a Qwen3.5-2B HybridDiffusion reproduction. The directory names, dataset_info.json, state.json, and Arrow shard boundaries are retained so the data can be downloaded and supplied to the existing training configuration without a lossy format conversion. Exact snapshot inventory Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arushhh/Llama-HybridDiffusion-processed-data-run1.text-generation1M<n<10M0 likes1k downloads1mo agoHugging Face13jeffry77 /Rule-VLN Rule-VLN Dataset Rule-VLN is a rule-compliant outdoor vision-and-language navigation benchmark built on the Touchdown / StreetLearn urban navigation environment. It studies whether navigation agents can follow language instructions while also complying with semantic traffic rules, such as regulatory signs that prohibit otherwise reachable movements. This dataset accompanies the paper: Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric… See the full description on the dataset page: https://huggingface.co/datasets/jeffry77/Rule-VLN.text-generation1 likes919 downloads3mo agoHugging Face14RussianNLP /russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills - detection of natural language inference, commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating models and an overall leaderboard of transformer models for the Russian language.text-classification100K<n<1M35 likes784 downloads3y agoHugging Face15ru-dataset /agent-think-tool_use Agent Think Tool Use Датасет многошаговых агентных сессий для дообучения моделей работе с кодом, инструментами и инженерными задачами. Записи содержат пользовательские требования, комментарии агента во время работы, decision summaries, вызовы инструментов, результаты запусков, обработку ошибок и финальную проверку. Каждый shard представляет отдельную связанную сессию, а не отдельный вопрос и ответ. Данные охватывают исследование задачи, работу с документацией, проектирование… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/agent-think-tool_use.tabulartext-generationn<1K2 likes744 downloads3d agoHugging Face16IlyaGusev /ru_turbo_saiga Saiga Dataset of ChatGPT-generated chats in Russian. Based on the Baize paper. Code: link. Prompt: Идёт диалог между пользователем и ИИ ассистентом. Пользователь и ассистент общаются на тему: {{seed}} Реплики человека начинаются с [Пользователь], реплики ассистента начинаются с [Ассистент]. Пользователь задаёт вопросы на основе темы и предыдущих сообщений. Пользователь обрывает беседу, когда у него не остается вопросов. Ассистент даёт максимально полные, информативные, точные и… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_saiga.text-generation10K<n<100K29 likes696 downloads3y agoHugging Face17deepvk /cultura_ru_edu Cultura-Ru-Edu The Cultura-Ru-Edu dataset consists of Russian educational web pages filtered from the uonlp/CulturaX dataset. The dataset creation was inspired by HuggingFaceFW/fineweb-edu, but with a focus on the Russian language. By filtering the dataset based on educational criteria, the Cultura-Ru-Edu dataset is both high-quality and large enough to train a Russian-focused language model for tasks requiring knowledge of the world. Dataset curation To create this… See the full description on the dataset page: https://huggingface.co/datasets/deepvk/cultura_ru_edu.texttext-generation100M<n<1B16 likes682 downloads2y agoHugging Face18Den4ikAI /russian_dialogues_2 Den4ikAI/russian_dialogues_2 Датасет русских диалогов для обучения диалоговых моделей. Количество диалогов - 1.6 миллиона Формат датасета: { 'sample': ['Привет', 'Привет', 'Как дела?'] } Citation: @MISC{russian_instructions, author = {Denis Petrov}, title = {Russian context dialogues dataset for conversational agents}, url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2}, year = 2023 } texttext-generation1M<n<10M18 likes633 downloads2y agoHugging Face19LLaMAX /BenchMAX_Rule-based Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios. We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.texttext-generation1K<n<10K1 likes603 downloads2y agoHugging Face20jablonkagroup /corral_runs_reports Corral – Evaluation Score Reports Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments. The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.tabulartext-generationn<1K0 likes541 downloads3mo agoHugging Face21ru-dataset /ru-dataset-small 🇷🇺 RU Dataset 1 Русскоязычный SFT-датасет для дообучения языковых моделей. Основной фокус — программирование, алгоритмы, архитектура ПО, математика и следование инструкциям. Все ответы развёрнутые, с reasoning-блоками <think> перед ответом. 🔄 Датасет активно пополняется. Новые диалоги добавляются регулярно — подпишитесь на обновления репозитория, чтобы не пропустить. Формат Стандартный chat-формат, совместимый с HuggingFace Datasets и большинством… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/ru-dataset-small.texttext-generation1K<n<10K2 likes513 downloads2mo agoHugging Face22lmqg /qg_ruquad[SberSQuAD](https://huggingface.co/datasets/sberquad) dataset for question generation (QG) task.text-generation10K<n<100K3 likes500 downloads4y agoHugging Face23cointegrated /ru-paraphrase-NMT-Leipzig Dataset Card for cointegrated/ru-paraphrase-NMT-Leipzig Dataset Summary The dataset contains 1 million Russian sentences and their automatically generated paraphrases. It was created by David Dale (@cointegrated) by translating the rus-ru_web-public_2019_1M corpus from the Leipzig collection into English and back into Russian. A fraction of the resulting paraphrases are invalid, and should be filtered out. The blogpost "Перефразирование русских текстов: корпуса, модели… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/ru-paraphrase-NMT-Leipzig.text-generation100K<n<1M12 likes483 downloads4y agoHugging Face24KirillR /Infinity-Instruct-RU-Synthetic Infinity-Instruct-RU-Synthetic A large-scale Russian-language instructional dataset based on Infinity-Instruct by BAAI. This is not a translation of English answers — it is an independent Russian-language dataset, where only the instructions are sourced from the original set, and all answers are newly generated in Russian from scratch. To translate the instructions, YandexGPT-5-Lite-8B-instruct was used with a specially fine-tuned LoRA adapter designed for this dataset. The original… See the full description on the dataset page: https://huggingface.co/datasets/KirillR/Infinity-Instruct-RU-Synthetic.texttext-generation1M<n<10M6 likes483 downloads1y agoHugging Face25d0rj /ru-instruct Карточка датасета Скомбинирован из нескольких популярных датасетов, переведённых автоматически. Отфильтрован на предмет артефактов перевода (спасибо модели Den4ikAI/nonsense_gibberish_detector). Дедуплицирован SimHash'ом. Обученной на нём модели пока не завёз, in progress. Состав Собрал из этих переведённых: d0rj/OpenOrca-ru (от Open-Orca/OpenOrca) d0rj/OpenHermes-2.5-ru (от teknium/OpenHermes-2.5) d0rj/dolphin-ru (от ehartford/dolphin) d0rj/alpaca-cleaned-ru (от… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/ru-instruct.texttext-generation100K<n<1M9 likes449 downloads2y agoHugging Face26RockMan256 /home_assistant_train_ru home_assistant_train_ru Русифицированная версия датасета acon96/Home-Assistant-Requests-V2 (41 798 примеров). Предназначена для дообучения маленьких моделей Home Assistant на русском языке (инструментальные вызовы, intent-классификация, управление устройствами). Что переведено User-запросы — полностью переведены на русский (ты-форма, неформально: «ты», не «вы»). Текстовые ответы assistant (естественный язык, идущий клиенту) — переведены. System-промпты, описания… See the full description on the dataset page: https://huggingface.co/datasets/RockMan256/home_assistant_train_ru.text-generation10K<n<100K0 likes446 downloads2mo agoHugging Face27sxiong /DHSA_RULER RULER Evaluation Data This dataset contains pre-generated JSONL files for the RULER long-context evaluation benchmark, used in Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference (ICML 2026 Spotlight). RULER is designed to evaluate effective context length and long-context behavior beyond simple retrieval, covering retrieval, multi-hop tracing, aggregation, and question answering style tasks. The files are organized by target… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/DHSA_RULER.text-generation10K<n<100K1 likes419 downloads2mo agoHugging Face28PotatoHD /ru-text-corpus Description 798k deduplicated Russian documents (1.6B tokens) from FineWeb-2 (rus_Cyrl), ru-StackOverflow, Pikabu, Habr, Russian Wikipedia and news. Markdown- and code-bearing (StackOverflow/Habr/Pikabu are text_markdown). Filtered to >=200 chars and >=30% Cyrillic. Columns: text, src. Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is. Usage from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/PotatoHD/ru-text-corpus.texttext-generation100K<n<1M0 likes390 downloads3mo agoHugging Face29VenusChenyy /RULER_50 RULER_50 Official-Code Qwen3 Subset This dataset is a fixed 50-sample-per-group subset of RULER synthetic tasks. It was generated from the official NVIDIA/RULER GitHub code, not from a third-party pre-generated mirror. Official generation source: Repository: https://github.com/NVIDIA/RULER Branch: main Commit: 38da79d79519ef87aa46ae804f838e1eab7f86d7 Generation entrypoint: scripts/data/prepare.py Benchmark config: scripts/synthetic.yaml Generation settings: tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/VenusChenyy/RULER_50.text-generation1 likes390 downloads1mo agoHugging Face30wwewtech /russian-it-community-corpus 📦 Russian IT Community Corpus (RICC) Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture. The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.tabulartext-generation1M<n<10M1 likes383 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.