CoolFace
Datasetpublic

davanstrien/paddleocr-vl15-final-test

Document Processing using PaddleOCR-VL-1.5 (OCR mode) This dataset contains OCR results from images in davanstrien/ufo-ColPali using PaddleOCR-VL-1.5, an ultra-compact 0.9B SOTA OCR model. Processing Details Source Dataset: davanstrien/ufo-ColPali Model: PaddlePaddle/PaddleOCR-VL-1.5 Task Mode: ocr - General text extraction to markdown format Number of Samples: 10 Processing Time: 8.3 min Processing Date: 2026-01-30 11:55 UTC Configuration Image… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/paddleocr-vl15-final-test.

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes6downloads
Dataset Card

Document Processing using PaddleOCR-VL-1.5 (OCR mode)

This dataset contains OCR results from images in davanstrien/ufo-ColPali using PaddleOCR-VL-1.5, an ultra-compact 0.9B SOTA OCR model.

Processing Details

Configuration

  • Image Column: image
  • Output Column: paddleocr_1.5_ocr
  • Dataset Split: train
  • Batch Size: 1
  • Smart Resize: Enabled
  • Max Output Tokens: 512
  • Backend: Transformers (batch inference)

Model Information

PaddleOCR-VL-1.5 is a state-of-the-art, resource-efficient model for document parsing:

  • 🎯 SOTA Performance - 94.5% on OmniDocBench v1.5
  • 🧩 Ultra-compact - Only 0.9B parameters
  • 📝 OCR mode - General text extraction
  • 📊 Table mode - HTML table recognition
  • 📐 Formula mode - LaTeX mathematical notation
  • 📈 Chart mode - Structured chart analysis
  • 🔍 Spotting mode - Text spotting with localization
  • 🔖 Seal mode - Seal and stamp recognition
  • 🌍 Multilingual - Support for multiple languages
  • Fast - Efficient batch inference

Task Modes

  • OCR: Extract text content to markdown format
  • Table Recognition: Extract tables to HTML format
  • Formula Recognition: Extract mathematical formulas to LaTeX
  • Chart Recognition: Analyze and describe charts/diagrams
  • Spotting: Text spotting with location information
  • Seal Recognition: Extract text from seals and stamps

Dataset Structure

The dataset contains all original columns plus:

  • paddleocr_1.5_ocr: The extracted content based on task mode
  • inference_info: JSON list tracking all OCR models applied to this dataset

Usage

python
from datasets import load_dataset
import json

# Load the dataset
dataset = load_dataset("{output_dataset_id}", split="train")

# Access the extracted content
for example in dataset:
    print(example["paddleocr_1.5_ocr"])
    break

# View all OCR models applied to this dataset
inference_info = json.loads(dataset[0]["inference_info"])
for info in inference_info:
    print(f"Task: {info['task_mode']} - Model: {info['model_id']}")

Reproduction

This dataset was generated using the uv-scripts/ocr PaddleOCR-VL-1.5 script:

bash
uv run https://huggingface.co/datasets/uv-scripts/ocr/raw/main/paddleocr-vl-1.5.py \
    davanstrien/ufo-ColPali \
    <output-dataset> \
    --task-mode ocr \
    --image-column image \
    --batch-size 1

Performance

  • Model Size: 0.9B parameters
  • Benchmark Score: 94.5% SOTA on OmniDocBench v1.5
  • Processing Speed: ~0.02 images/second
  • Backend: Transformers batch inference

Generated with 🤖 UV Scripts