datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
russian-old-orthography-ocr
Basic Description
Dataset contains source images and human-readable extracted texts. All texts were published in Russia in the 19th century and written using pre-reform orthography.
The dataset is designed to train and evaluate optical character recognition systems for texts published in Russian before the orthographic reform (1917).
Data structure
For each text there is a file with its image and the text corresponding to this image. The names of these files are the same… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-old-orthography-ocr.russian-road-signs
Датасет размеченных знаков
Датасет размеченных дорожных знаков для задач компьютерного зрения и детекции объектов.
Загрузка
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="Dognellaf/russian-road-signs",
repo_type="dataset",
local_dir="./russian-road-signs"
)
Описание
Датасет содержит размеченные вручную кадры из видеозаписей с российскими дорожными знаками. Разметка в формате YOLO.
Изображений: 43 851 (JPEG)… See the full description on the dataset page: https://huggingface.co/datasets/Dognellaf/russian-road-signs.SpatialBenchSpatialBench evaluates model performance on spatial understanding. We design positional, existence, counting, reaching and size comparasion tasks.
In this HF dataset, SpatialBench RGB & Depth images, questions, answers and meta data are provided.
Paper:
https://arxiv.org/abs/2406.13642
GitHub repo:
https://github.com/BAAI-DCAI/SpatialBot
SpatialBot, a VLM with precise depth understanding:
https://huggingface.co/RussRobin/SpatialBot
russian-handwriting-ocr
Russian Handwritten Text Recognition Dataset
Датасет для распознавания русских рукописных текстов (сочинений).
Описание
Этот датасет содержит изображения рукописных русских текстов с их расшифровкой.
Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста.
Статистика
Всего образцов: 13050
Train: 11745
Validation: 1305
Уникальных текстов: 575
Средняя длина текста: 3790 символов
Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/rustensai/russian-handwriting-ocr.libero_unlearned_orange_juiceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1693,
"total_frames": 273465,
"total_tasks": 40,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10.0,
"splits": {
"train": "0:1693"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/leonardo-russo/libero_unlearned_orange_juice.russian-handwriting-ocr
Russian Handwritten Text Recognition Dataset
Датасет для распознавания русских рукописных текстов (сочинений).
Описание
Этот датасет содержит изображения рукописных русских текстов с их расшифровкой.
Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста.
Статистика
Всего образцов: 13050
Train: 11745
Validation: 1305
Уникальных текстов: 575
Средняя длина текста: 3790 символов
Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/Limerencii/russian-handwriting-ocr.flux-50k
FLUX-Generated Synthetic Dataset from BLIP3o Captions
This is a synthetic dataset where all images were generated using the FLUX text-to-image model from captions in the russwang/blip3o-long-caption-50k dataset.
Dataset Structure
Data Fields
id: Unique identifier (global index)
caption: Text prompt used to generate the image
image: Generated image (PIL Image)
caption_length: Length of caption in characters
generation_time: Time taken to generate image… See the full description on the dataset page: https://huggingface.co/datasets/russwang/flux-50k.russian_docs_bypageVDDVDD: Varied Drone Dataset for Semantic Segmentation
Published in Journal of Visual Communication and Image Representation.
Paper: https://arxiv.org/abs/2305.13608
GitHub Repo: https://github.com/RussRobin/VDD
This HF repo contains VDD source images and annotations.
If you would like to use Integrated Drone Dataset (IDD), which is a combination of VDD, UDD and UAVid, please refer to: https://huggingface.co/datasets/RussRobin/IDD
SpatialQASpatialQA enhances the model's spatial understanding capabilities by helping it comprehend and utilize depth maps.
In this HF dataset, SpatialQA.json and high-level images are provided. Please also download images in Bunny_695k for low and middle-level images.
How to use this dataset
Download images and json from this repo
Download Bunny_695k
Prepare depth map for coco_2017 and visual_genome. Please refer to our instructions.
File structure:
/images
/images/coco_2017… See the full description on the dataset page: https://huggingface.co/datasets/RussRobin/SpatialQA.RussianVibe-data12_popular_russia_mushrooms_edible_poisonousrussian_ocr_small
Russian OCR Small Dataset
Combined dataset for Russian text recognition (OCR).
Source Datasets
adasdaadadad/Car_plate_OCR_dataset
constantinwerner/cyrillic-handwriting-dataset
nvidia/OCR-Synthetic-Multilingual-v1
Structure
car_plate: License plate images (1,500 samples)
handwriting: Handwritten words/phrases (1,544 samples)
printed: Synthetic printed text (1,000 samples)
License
Dataset Governing Terms: Use of the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Foximaz/russian_ocr_small.Louis.Vuitton.Product.prices.Russia
Louis Vuitton web scraped data
About the website
The luxury fashion industry in the EMEA region, particularly in Russia, is characterized by a growing demand for high-end products from renowned brands. Louis Vuitton, a global leader in this industry, caters to this escalating demand through their extensive range of luxury clothing, accessories, and luggage. The brand has significantly increased its presence in Russia by leveraging the power of Ecommerce, effectively… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Louis.Vuitton.Product.prices.Russia.yolo5_russianlicenseplates_detectMMstar_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the MMStar dataset.
MMStar is an elite vision-language benchmark consisting of high-quality, non-redundant multimodal samples designed to evaluate 6 core capabilities and 18 detailed dimensions of Large Multimodal Models (LMMs). Kazakh and Russian versions serve as a benchmark for evaluating how well models perform advanced visual recognition and multi-step reasoning while avoiding "data leakage"… See the full description on the dataset page: https://huggingface.co/datasets/issai/MMstar_Kazakh_Russian.SpatialQA-ESpatialQA-E is a robot manipulation dataset focusing on spatial relationship understanding.
Paper:
https://arxiv.org/abs/2406.13642
GitHub repo:
https://github.com/BAAI-DCAI/SpatialBot
SpatialBot-general QA, a VLM with precise depth understanding:
https://huggingface.co/RussRobin/SpatialBot
SpatialBench, the spatial understanding benchmark in general QA:
https://huggingface.co/datasets/RussRobin/SpatialBench
RealWorldQA_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the RealWorldQA dataset. As a multimodal benchmark, RealWorldQA is specifically designed to evaluate the real-world spatial understanding and visual reasoning of models. Kazakh and Russian versions serve as a benchmark for evaluating how well models can understand physical environments, spatial relationships, and object attributes based on real-world images when prompted in Kazakh or Russian.… See the full description on the dataset page: https://huggingface.co/datasets/issai/RealWorldQA_Kazakh_Russian.cerno_dataset_yoloNet.a.Porter.Product.prices.Russia
Net-a-Porter web scraped data
About the website
The EMEA fashion industry, particularly in Russia, has been experiencing substantial growth in online channels due to increased internet penetration and smartphone usage. A significant player in this advancement is Net-a-Porter. This platform belongs to the luxury ecommerce industry, offering a wide range of premium brands. With the shift towards digital platforms in the shopping behavior of consumers, Net-a-porter is making… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Net.a.Porter.Product.prices.Russia.imnet1k_borzoi_Russian_wolfhoundrussoeng-russian-paintings-t2i-last-1000blip3o-long-caption-50k
BLIP3o Long Caption Dataset (50K Subset)
This dataset is a subset of the BLIP3o/BLIP3o-Pretrain-Long-Caption dataset, containing the first 50,000 samples.
Statistics
Total samples: 50,000
Average caption length: 624.4 characters
Min caption length: 125 characters
Max caption length: 983 characters
Dataset Structure
Data Fields
id: Unique identifier for each sample
caption: Long-form image caption
caption_length: Length of the caption in characters… See the full description on the dataset page: https://huggingface.co/datasets/russwang/blip3o-long-caption-50k.russian_memes
Russian Meme Dataset
This dataset contains annotated Russian-language image memes. It is designed for evaluating multimodal and text-based models on meme understanding, with a focus on toxicity, sarcasm, irony, and local cultural references.
The dataset is released in two CSV versions:
dataset_basic.csv — a compact annotation table with image filenames and target labels.
dataset_extended.csv — an extended table that also includes OCR text, visual scene descriptions, and… See the full description on the dataset page: https://huggingface.co/datasets/SoulQrat/russian_memes.StitchBenchThis is the official HF repo for StitchBench, a comprehensive image stitching benchmark, proposed in Object-level Geometric Structure Preserving for Natural Image Stitching (OBJ-GSP).
You can download the images by:
from huggingface_hub import snapshot_download
repo_id = "RussRobin/StitchBench"
local_dir = "./StitchBench"
snapshot_download(repo_id=repo_id, repo_type="dataset", local_dir=local_dir)
AAAI 2025 Paper:
https://arxiv.org/abs/2402.12677
GitHub:
https://github.com/RussRobin/OBJ-GSP… See the full description on the dataset page: https://huggingface.co/datasets/RussRobin/StitchBench.MedHallTuneeng-russian-paintings-t2i-lastoptic_QA_pdf_russianMr.Porter.Product.prices.Russia
Mr Porter web scraped data
About the website
Mr Porter is a prominent operator in the online retail industry within the EMEA region, specifically in Russia. The E-commerce industry in Russia is growing rapidly, amidst the increasing tech-savviness and online shopping habits of consumers. Mr Porters foundation on sophisticated technology offers them a competitive edge, especially in tailoring to Russias vast and diverse consumer base. E-commerce, particularly online… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Mr.Porter.Product.prices.Russia.
