CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01p4ulbr4dl3y /glm-ocr-ru-dataset GLM-OCR Russian Dataset (Interfax) This dataset was generated programmatically for fine-tuning Vision-Language models, specifically GLM-OCR or Qwen-VL / Qwen2-VL, on Russian document OCR. Dataset Details Source: News texts from the Russian news agency "Interfax". Size: 4,998 samples (multi-paragraph image-text pairs). Format: ShareGPT VLM format (compatible with LLaMA-Factory out-of-the-box). Image format: WebP (quality 85) for minimal disk space footprint (~150… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/glm-ocr-ru-dataset.image1K<n<10K0 likes357 downloads3mo agoHugging Face02p4ulk /ISP-AD ISP-AD: The Industrial Screen Printing Anomaly Detection Dataset The ISP-AD Dataset is a large-scale industrial visual anomaly detection benchmark designed for unsupervised, self-supervised, and supervised learning. It features subtle, weakly contrasted surface defects embedded within structured screen-printed patterns with high permitted design variability. Comprising 559,049 samples, ISP-AD is one of the largest publicly available industrial anomaly detection datasets to date… See the full description on the dataset page: https://huggingface.co/datasets/p4ulk/ISP-AD.imageimage-classificationn<1K0 likes121 downloads24d agoHugging Face03RyanWy /P4NSU P4NSU: Projection-based Pretraining for Nonlinear Sparse Unmixing in Spectral Imaging The main codes of P4NSU can be seen in Github. 🚀 How to Use You can download this dataset directly in your Python script: !pip install huggingface_hub -q from huggingface_hub import snapshot_download # Download dataset to a local folder called 'P4NSU' snapshot_download(repo_id="RyanWy/P4NSU", repo_type="dataset", local_dir="./P4NSU") # Then… See the full description on the dataset page: https://huggingface.co/datasets/RyanWy/P4NSU.imagen<1K0 likes107 downloads8mo agoHugging Face04jablonkagroup /cyp_p450_3a4_inhibition_veith_et_al-multimodalimage100K<n<1M0 likes84 downloads1y agoHugging Face05jablonkagroup /cyp_p450_2d6_inhibition_veith_et_al-multimodalimage100K<n<1M0 likes55 downloads1y agoHugging Face06p4ulbr4dl3y /qwen-chart-dataset-v2 Qwen Chart Dataset v2 A multimodal dataset designed for training Vision-Language Models (VLMs) to analyze and interpret charts and graphs. Dataset Structure Format: Image-Text pairs with detailed descriptions. Samples: 1,000+ charts covering 10+ distinct types. Chart Types: Bar, line, scatter, pie, heatmap, box, radar, and more. Use Case Specifically optimized for models to extract trends, exact values, and statistical correlations from visual… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/qwen-chart-dataset-v2.imagevisual-question-answering1K<n<10K0 likes54 downloads6mo agoHugging Face07jablonkagroup /cyp_p450_2c9_inhibition_veith_et_al-multimodalimage100K<n<1M0 likes42 downloads1y agoHugging Face08andlyu /pack4_v6_p4imagen<1K0 likes36 downloads9mo agoHugging Face09jablonkagroup /cyp_p450_2c19_inhibition_veith_et_al-multimodalimage100K<n<1M0 likes32 downloads1y agoHugging Face10p4ulbr4dl3y /doc2json-vlm-full Doc2JSON VLM Full Синтетический датасет изображений российских документов для fine-tuning VLM, извлекающей данные в JSON по динамически заданной схеме. Опубликован 4 августа 2026 года. Содержит 3 280 примеров из 700 исходных документов. Типы документов Паспорт РФ. Счет на оплату. Акт выполненных работ. УПД. Договор поставки. Договор оказания услуг. Договор подряда. Договор аренды. Лицензионный договор. Спецификация. Дополнительное соглашение. Акт сверки. ТОРГ-12.… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/doc2json-vlm-full.imageimage-to-text1K<n<10K0 likes29 downloads2mo agoHugging Face11SprintML /P4Ms-hackathon-vision-taskimage1K<n<10K0 likes24 downloads5mo agoHugging Face12p4ulbr4dl3y /document-ocr-vlm-dataset-5kimage1K<n<10K0 likes24 downloads3mo agoHugging Face13p4ulbr4dl3y /passport-ocr-vlmimagen<1K0 likes23 downloads2mo agoHugging Face14p4ulbr4dl3y /dogovors-unlimited-ocr Dogovors Unlimited-OCR Dataset OCR/document-layout dataset prepared for fine-tuning baidu/Unlimited-OCR. Files train.jsonl contains one JSON object per document. images/ contains the page images referenced by relative path. JSONL Schema { "images": [ "images/doc_001_page_001.jpg", "images/doc_001_page_002.jpg" ], "question": "Multi page parsing.", "answer": "<PAGE><|det|>title [400, 60, 660, 73]<|/det|>... <PAGE><|det|>text [100… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/dogovors-unlimited-ocr.imagen<1K0 likes22 downloads3mo agoHugging Face15p4ulbr4dl3y /multipage-contractsimagen<1K0 likes22 downloads2mo agoHugging Face16crazylearners /p4demoimagen<1K0 likes20 downloads3y agoHugging Face17P4rz1val /SyntheticBeverages Med-10 Synthetic Beverage Dataset This dataset is part of the Med-10 project, aimed at evaluating how different degradation types affect object detection performance. It contains procedurally generated synthetic images of beverage bottles rendered under varying conditions. 🧪 Dataset Variants The dataset consists of 19 different variants, each applying one type of degradation to isolate its effect. Variants are grouped into two main categories: Wild: Fully randomized… See the full description on the dataset page: https://huggingface.co/datasets/P4rz1val/SyntheticBeverages.imageobject-detection1K<n<10K0 likes17 downloads1y agoHugging Face18kavinh07 /ocr_dataset_shamadhan_synth_30k_p4 NID OCR Extended Dataset Built from kavinh07/ocr_dataset_shamadhan_synth_30k_p2 with 9,600 additional confusion-pair training images. Split Samples train 155,400 validation 35,363 Sources kavinh07/ocr_dataset_shamadhan_synth_30k_p2 — original synthetic + shamadhan real data synthetic_confusion — 9,600 confusion-pair images (train only) Columns Column Type Description image Image Cropped NID field (RGB) text string… See the full description on the dataset page: https://huggingface.co/datasets/kavinh07/ocr_dataset_shamadhan_synth_30k_p4.imageimage-to-text100K<n<1M0 likes10 downloads3mo agoHugging Face19jablonkagroup /cyp_p450_1a2_inhibition_veith_et_al-multimodalimage100K<n<1M0 likes8 downloads1y agoHugging Face20p4ulbr4dl3y /contracts-ocr-1kimagen<1K0 likes7 downloads3mo agoHugging Face21daaxila /twitter-TianxinKitten-2025.09.04-1963616789667201031-wCBahDZwB4h_p4o-part1imagen<1K0 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.