CoolFace
Modelpublic

martinbadrous/YOLOv8-Historical-Document-Detection

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes
Model Card

πŸ›οΈ YOLOv8 β€” Historical Document Ornament Detector

Automatic detection of typographic ornaments in 16th–18th century printed documents. Developed as part of the TypoRef project at PolyTech Tours.

πŸ“Œ Model Summary

PropertyDetails
πŸ—οΈ ArchitectureYOLOv8 (Ultralytics)
🎯 TaskObject Detection
πŸ“Š mAP@5095%
πŸ—‚οΈ Dataset50+ expert-annotated historical document pages
πŸ“… Document period16th – 18th century printed books
βš™οΈ FrameworkPyTorch + Ultralytics
πŸ“‰ Processing speedup20% faster than manual workflow
πŸ“œ LicenseMIT

🧠 What This Model Does

This model detects and localizes typographic ornaments and decorative graphic elements in scanned pages of early modern European printed books.

It was built to replace a slow, fully manual cataloguing process for the TypoRef digital humanities project, enabling automated analysis of thousands of document pages that would otherwise require extensive expert annotation.

Detected classes: typographic ornaments, decorative initials, vignettes, and other graphic elements typical of 16th–18th century printing.


πŸ“ˆ Performance

MetricScore
mAP@5095%
Training duration6 months iterative refinement
Annotations integrated50+ pages in 2 months
Processing time reduction20% vs previous pipeline

πŸš€ How to Use

python
from ultralytics import YOLO

# Load the model
model = YOLO("best.pt")

# Run inference on a document scan
results = model("your_document_scan.jpg", conf=0.35)

# Show results
results[0].show()

# Save annotated image
results[0].save("output.jpg")

πŸ—‚οΈ Training Data

  • β€”Source: Historical printed books from the TypoRef corpus (16th–18th century)
  • β€”Annotations: Expert-annotated by digital humanities researchers at PolyTech Tours
  • β€”Volume: 50+ annotated document pages
  • β€”Augmentation: Standard YOLOv8 augmentation pipeline

⚠️ Limitations

  • β€”Optimized for black-and-white or greyscale document scans
  • β€”Performance may degrade on very low-resolution scans (< 150 DPI)
  • β€”Trained on Western European printing conventions β€” may generalize poorly to other traditions

πŸ”— Related Resources


πŸ‘€ Author

Martin Badrous β€” Computer Vision & Deep Learning Engineer

![LinkedIn](https://www.linkedin.com/in/martinbadrous) ![GitHub](https://github.com/martinbadrous) ![HuggingFace](https://huggingface.co/martinbadrous)