datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLaVA-OneVision-Data-ru
LLaVA-OneVision-Data-ru
Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate.
Almost all datasets have been translated, except for the following:
["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"]
Usage
import datasets
data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.ttrpg-rpg-fandom-com-ru
ttrpg-rpg-fandom-com-ru (RPG Fandom RU Dataset)
[Russian version below / Русская версия ниже]
Описание
Этот датасет содержит полный дамп вики rpg.fandom.com/ru/, конвертированный в чистый Markdown. Он предназначен для использования в системах RAG (Retrieval-Augmented Generation), дообучения языковых моделей (LLM) и исследований в области НРИ.
Создатель и контакты
Вебсайт: exnihilum.info
GitHub: exnpub
Репозиторий проекта: dataset-rpg-fandom-com-ru… See the full description on the dataset page: https://huggingface.co/datasets/exnihilum/ttrpg-rpg-fandom-com-ru.ru_paradetox
ParaDetox: Text Detoxification with Parallel Data (Russian)
This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit
[2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.yeji-bazi-rules
██████╗ █████╗ ███████╗██╗ ██████╗ ██╗ ██╗██╗ ███████╗███████╗
██╔══██╗██╔══██╗╚══███╔╝██║ ██╔══██╗██║ ██║██║ ██╔════╝██╔════╝
██████╔╝███████║ ███╔╝ ██║ ██████╔╝██║ ██║██║ █████╗ ███████╗
██╔══██╗██╔══██║ ███╔╝ ██║ ██╔══██╗██║ ██║██║ ██╔══╝ ╚════██║
██████╔╝██║ ██║███████╗██║ ██║ ██║╚██████╔╝███████╗███████╗███████║
╚═════╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝╚══════╝╚══════╝
⚡ INTERPRETATION RULEBOOK ⚡
> ACCESS… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-bazi-rules.eleusis-calibrated-rules
Eleusis Calibrated Rules — 100-turn reward calibration
A calibrated rule dataset for the single-player Eleusis inductive-reasoning
environment. It extends the 26-rule Hugging Face benchmark with controlled
static, transition, conditional, periodic, chunk, higher-order history, global
history, and compositional rule families.
Source benchmark: Hugging Face Eleusis.
Dataset version: v2.1-frontier-calibrated-100turn-20260812Protocol: eleusis-100-v11
The structural, GPT Sol… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-calibrated-rules.CodeAnything-1.835M
CodeAnything 1.835M SFT and evaluation release
Gated public release containing the clean 1,835,476-sample training set, the
800-sample/16-domain evaluation set, paper-model predictions and rendered
outputs, and raw per-sample rating records.
Layout
training/
manifest/
all_training_v5.jsonl
all_training_v5.jsonl.idx
all_training_v5.jsonl.true_lengths.u32
shards/<domain>/ exact media/code closure (tar shards)
evaluation/
benchmark/… See the full description on the dataset page: https://huggingface.co/datasets/Ruler138/CodeAnything-1.835M.DALL-E-Prompts-OpenAI-ChatGPT
Dataset Card for Dataset Name
Dataset Summary
This dataset has been generated using Prompt Generator for OpenAI's DALL-E.
Languages
English
Dataset Structure
1.000.000 Prompts
minecraft-crafting-vqa-ru-en
Minecraft Crafting & Gameplay VQA Dataset (RU/EN)
A bilingual, multimodal dataset specifically designed for fine-tuning Vision-Language Models (VLMs) such as Qwen2.5-VL and Qwen3-VL. This dataset is optimized for training AI models to recognize crafting recipes, understand in-game user interfaces, and trigger function calls (Tool Calling).
Repository Structure
The dataset is cleanly separated into language locales. Each folder contains its own data.json annotation file… See the full description on the dataset page: https://huggingface.co/datasets/KuroTo4ka/minecraft-crafting-vqa-ru-en.geometry-reasoning-ru
Задачи по геометрии МЦНМО
Описание
В датасете собраны задачи, решения, подсказки и комментарии по геометрии различного уровня сложности на русском языке. Для некоторых задач добавлены картинки.
Датасет может применяться для оценки способности LLM строить причинно-следственные связи и делать правильные выводы на русском языке.
Multi-view-CXR
EVOKE
Patient-Specific Multimodal Learning with Multi-View Contrastive Alignment for Chest X-ray Report Generation
Radiology reports are crucial for planning treatment strategies and facilitating effective doctor-patient communication. However, the manual creation of these reports places a significant burden on radiologists. While automatic radiology report generation presents a promising solution, existing methods often rely on single-view radiographs, which constrain… See the full description on the dataset page: https://huggingface.co/datasets/MK-runner/Multi-view-CXR.RealWorldQA_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the RealWorldQA dataset. As a multimodal benchmark, RealWorldQA is specifically designed to evaluate the real-world spatial understanding and visual reasoning of models. Kazakh and Russian versions serve as a benchmark for evaluating how well models can understand physical environments, spatial relationships, and object attributes based on real-world images when prompted in Kazakh or Russian.… See the full description on the dataset page: https://huggingface.co/datasets/issai/RealWorldQA_Kazakh_Russian.svg-multimodal-rubrics
SVG Multimodal Rubrics
A multimodal dataset of SVG code generation samples with natural language descriptions and evaluation rubrics. Each sample pairs a detailed prompt (Markdown) with its corresponding SVG source code, covering animations, 3D scenes, games, and visual effects.
Designed for training and evaluating models on visual code generation — generating complex, interactive SVG artwork from natural language descriptions.
Overview
Item
Details
Samples… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/svg-multimodal-rubrics.SceneChain-12K
SceneChain-12K
SceneChain-12K is a multi-turn scene editing conversation dataset for training vision-language models to generate and iteratively refine 3D indoor scenes.
Data Format
Each sample is a JSON object with:
messages: Multi-turn conversation following OpenAI chat format (system/user/assistant)
images: List of rendered scene image paths (relative to dataset root)
Conversation Structure
System: Scene editing instructions and tool definitions
User:… See the full description on the dataset page: https://huggingface.co/datasets/runder1/SceneChain-12K.
