datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glm-ocr-ru-dataset
GLM-OCR Russian Dataset (Interfax)
This dataset was generated programmatically for fine-tuning Vision-Language models, specifically GLM-OCR or Qwen-VL / Qwen2-VL, on Russian document OCR.
Dataset Details
Source: News texts from the Russian news agency "Interfax".
Size: 4,998 samples (multi-paragraph image-text pairs).
Format: ShareGPT VLM format (compatible with LLaMA-Factory out-of-the-box).
Image format: WebP (quality 85) for minimal disk space footprint (~150… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/glm-ocr-ru-dataset.ISP-AD
ISP-AD: The Industrial Screen Printing Anomaly Detection Dataset
The ISP-AD Dataset is a large-scale industrial visual anomaly detection benchmark designed for unsupervised, self-supervised, and supervised learning. It features subtle, weakly contrasted surface defects embedded within structured screen-printed patterns with high permitted design variability.
Comprising 559,049 samples, ISP-AD is one of the largest publicly available industrial anomaly detection datasets to date… See the full description on the dataset page: https://huggingface.co/datasets/p4ulk/ISP-AD.P4NSU
P4NSU: Projection-based Pretraining for Nonlinear Sparse Unmixing in Spectral Imaging
The main codes of P4NSU can be seen in Github.
🚀 How to Use
You can download this dataset directly in your Python script:
!pip install huggingface_hub -q
from huggingface_hub import snapshot_download
# Download dataset to a local folder called 'P4NSU'
snapshot_download(repo_id="RyanWy/P4NSU",
repo_type="dataset",
local_dir="./P4NSU")
# Then… See the full description on the dataset page: https://huggingface.co/datasets/RyanWy/P4NSU.cyp_p450_3a4_inhibition_veith_et_al-multimodalcyp_p450_2d6_inhibition_veith_et_al-multimodalqwen-chart-dataset-v2
Qwen Chart Dataset v2
A multimodal dataset designed for training Vision-Language Models (VLMs) to analyze and interpret charts and graphs.
Dataset Structure
Format: Image-Text pairs with detailed descriptions.
Samples: 1,000+ charts covering 10+ distinct types.
Chart Types: Bar, line, scatter, pie, heatmap, box, radar, and more.
Use Case
Specifically optimized for models to extract trends, exact values, and statistical correlations from visual… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/qwen-chart-dataset-v2.cyp_p450_2c9_inhibition_veith_et_al-multimodalpack4_v6_p4cyp_p450_2c19_inhibition_veith_et_al-multimodaldoc2json-vlm-full
Doc2JSON VLM Full
Синтетический датасет изображений российских документов для fine-tuning VLM,
извлекающей данные в JSON по динамически заданной схеме.
Опубликован 4 августа 2026 года. Содержит 3 280 примеров из 700 исходных
документов.
Типы документов
Паспорт РФ.
Счет на оплату.
Акт выполненных работ.
УПД.
Договор поставки.
Договор оказания услуг.
Договор подряда.
Договор аренды.
Лицензионный договор.
Спецификация.
Дополнительное соглашение.
Акт сверки.
ТОРГ-12.… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/doc2json-vlm-full.P4Ms-hackathon-vision-taskdocument-ocr-vlm-dataset-5kpassport-ocr-vlmdogovors-unlimited-ocr
Dogovors Unlimited-OCR Dataset
OCR/document-layout dataset prepared for fine-tuning baidu/Unlimited-OCR.
Files
train.jsonl contains one JSON object per document.
images/ contains the page images referenced by relative path.
JSONL Schema
{
"images": [
"images/doc_001_page_001.jpg",
"images/doc_001_page_002.jpg"
],
"question": "Multi page parsing.",
"answer": "<PAGE><|det|>title [400, 60, 660, 73]<|/det|>...
<PAGE><|det|>text [100… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/dogovors-unlimited-ocr.multipage-contractsp4demoSyntheticBeverages
Med-10 Synthetic Beverage Dataset
This dataset is part of the Med-10 project, aimed at evaluating how different degradation types affect object detection performance. It contains procedurally generated synthetic images of beverage bottles rendered under varying conditions.
🧪 Dataset Variants
The dataset consists of 19 different variants, each applying one type of degradation to isolate its effect. Variants are grouped into two main categories:
Wild: Fully randomized… See the full description on the dataset page: https://huggingface.co/datasets/P4rz1val/SyntheticBeverages.ocr_dataset_shamadhan_synth_30k_p4
NID OCR Extended Dataset
Built from kavinh07/ocr_dataset_shamadhan_synth_30k_p2
with 9,600 additional confusion-pair training images.
Split
Samples
train
155,400
validation
35,363
Sources
kavinh07/ocr_dataset_shamadhan_synth_30k_p2 — original synthetic + shamadhan real data
synthetic_confusion — 9,600 confusion-pair images (train only)
Columns
Column
Type
Description
image
Image
Cropped NID field (RGB)
text
string… See the full description on the dataset page: https://huggingface.co/datasets/kavinh07/ocr_dataset_shamadhan_synth_30k_p4.cyp_p450_1a2_inhibition_veith_et_al-multimodalcontracts-ocr-1ktwitter-TianxinKitten-2025.09.04-1963616789667201031-wCBahDZwB4h_p4o-part1
