CoolFace
Datasetpublic

small-models-for-glam/index-card-detection-v3

Dataset Card for Archival Index Card Detection — mixed collections A training dataset for object detection of index cards in archival scans. Combines four publicly-released collections — NLS Advocates Library single-card pages, US Navy Nurse Corps multi-card biographical sheets, Boston Public Library catalog cards, and Duke Rubenstein manuscript catalog cards — into a single object-detection schema. Dataset Details Dataset Description 1,425 archival… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v3.

sourceHugging Facecc0-1.0updated 4mo agoView on Hugging Face
0likes35downloads
Dataset Card

Dataset Card for Archival Index Card Detection — mixed collections

A training dataset for object detection of index cards in archival scans. Combines four publicly-released collections — NLS Advocates Library single-card pages, US Navy Nurse Corps multi-card biographical sheets, Boston Public Library catalog cards, and Duke Rubenstein manuscript catalog cards — into a single object-detection schema.

Dataset Details

Dataset Description

1,425 archival scans across four institutional collections. Constructed to extend the existing single-card NLS detector (`NationalLibraryOfScotland/archival-index-card-detector`) to handle multi-card scans and pre-cropped single-card images from other institutions, without regressing on the original NLS distribution.

  • Curated by: Daniel van Strien, Machine Learning Librarian, Hugging Face
  • Language(s): en (card content; English-only)
  • License: Mixed — see Source Data table below. Effectively unrestricted re-use for academic / research / non-commercial purposes; check the source dataset for commercial reuse.

Dataset Sources

  • Repository: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v3
  • Trained model: https://huggingface.co/small-models-for-glam/index-card-detector-v4
  • Demo Space: https://huggingface.co/spaces/small-models-for-glam/index-card-detector

Uses

Direct Use

Train an object detector to find archival index cards on:

  • Full archival page scans with one or more cards (NLS-style multi-card pages)
  • Multi-card sheets (2–9 cards arranged on a backing sheet, e.g. Navy biographical records)
  • Pre-cropped single-card images (BPL / Rubenstein style)

The accompanying YOLO26n checkpoint at small-models-for-glam/index-card-detector-v4 is fine-tuned from the NLS v1 baseline on this data.

Out-of-Scope Use

  • Not for OCR. Detection only — bbox tells you where the card is, not what it says. Pair with a downstream OCR/VLM model (NuExtract3, Qwen-VL, etc.).
  • Not for content classification. Single class card — does not distinguish card type, content, or blank/content.
  • English-only. Cards in other languages may not generalise.
  • Not a benchmark dataset. Train/val split is automatic and stratified per collection; not held back as a community evaluation set.

Dataset Structure

One row per image:

ColumnTypeDescription
imageImageRGB scan, original resolution
source_collectionstringOne of nls-advocates, navy-nurse-corps, bpl, rubenstein
source_repostringSource HF dataset id
source_row_idint64Index in the source dataset
source_urlstringLink to original item (when present in source)
textstringOCR or filename from source (preserved as-is)
objects.bboxlist[[float, 4]]xywh in original pixel coordinates
objects.categoryClassLabel(["card"])Single class; always 0
objects.box_sourcelist[string]Provenance per box: nls-original \sam3+human \auto-full

Per-collection breakdown:

CollectionRowsBoxesAvg boxes/rowAvg image size
nls-advocates100~1551.55 (incl. some has_card=False)~2900×1800
navy-nurse-corps25~853.4varied (3K×4K to 10K×6K)
bpl8008001.0~1175×710
rubenstein5005001.0~1555×1100
Total1,425~1,540

Dataset Creation

Curation Rationale

The existing NLS detector (v1) was trained on 100 NLS Advocates Library pages and reached 99.1% mAP@50:95 on that distribution, but had not been tested on multi-card scans or cards from other institutions. This dataset extends v1's training distribution while preserving the NLS samples to prevent catastrophic forgetting.

Source Data

Data Collection and Processing

Four source collections, all publicly hosted on the Hugging Face Hub:

CollectionSourcen usedLabel provenanceLicense (source)
NLS Advocates LibraryNationalLibraryOfScotland/nls-index-cards-object-detection100 (all)Original NLS bboxes (10 has_card=False negatives included)Per source repo
US Navy Nurse Corpsbiglam/index-cards-navy-nurse-corps25 (all)SAM3 bootstrap + human review/correctionPublic domain (US Govt)
Boston Public Librarybiglam/bpl-card-catalog800 (first rows)bbox = whole image (auto, pre-cropped)CC0-1.0
Duke Rubensteinbiglam/rubenstein-manuscript-catalog500 (first rows)bbox = whole image (auto, pre-cropped)CC0-1.0

Sloane Catalogues (biglam/sloane-catalogues) was considered and excluded — those scans are two-page manuscript catalog book spreads, not index cards.

Who are the source data producers?

Original archival cards were created by librarians and archivists at the four source institutions over the late 19th / 20th century. Digitisation by each institution and/or via Internet Archive.

Annotations

Annotation process

Three different paths into the same object-detection schema:

  1. 1.NLS Advocates (100 imgs): original bboxes from the NLS detector v1 training set, passed through unchanged. Includes 10 has_card=False pages as negatives.
  1. 1.Navy Nurse Corps (25 imgs): bootstrap with Meta's SAM3 (`uv-scripts/sam3`, --class-name "card" --confidence-threshold 0.15 on an A100). Outputs were reviewed and corrected in a custom single-page HTML bbox editor. One systematic SAM3 failure (an "envelope" box wrapping all cards on tall multi-card scans, ~95–98% image area) was auto-filtered before review. Final per-image: ~3.4 cards on average.
  1. 1.BPL + Rubenstein (1,300 imgs): auto-labelled bbox = [1, 1, W-2, H-2] based on the structural fact that source images are pre-cropped single cards filling the frame. No human review pass.

The objects.box_source column records which path produced each box.

Who are the annotators?
  • NLS: NLS internal team
  • Navy: Daniel van Strien (review and correction over SAM3 bootstrap)
  • BPL + Rubenstein: rule-based (no human pass)
Personal and Sensitive Information

Cards contain names, addresses, and biographical details of individuals (Navy Nurse Corps records in particular). All source datasets are from publicly-released collections — personal information is in the public archival record by virtue of original institutional release. Downstream users should respect each source institution's usage policies.

Bias, Risks, and Limitations

  • English-only. Cards in other languages are not represented.
  • US/UK archival conventions only. Card stocks, layouts, and aesthetics outside that tradition (e.g. continental European library catalogs, East Asian indices, hand-coloured cards) are out-of-distribution.
  • Multi-card variety concentrated in 25 navy images. Multi-card scans with very different layouts (e.g., card grids of 12+, severely overlapping cards) may not generalise.
  • No true negative non-card examples beyond the 10 NLS background pages. Newspaper clippings, photographs, and book pages are not in the training set; models trained on this may over-predict on those inputs. Pair with downstream filtering or train v5 with additional negatives.
  • BPL + Rubenstein labels are rule-based, not human-verified. bbox = whole image is a reasonable approximation for pre-cropped cards but slight cropping noise around the edges is not captured. Tight pixel-accuracy downstream tasks may want a hand-labelled refresh.

Recommendations

  • For tight cropping or pixel-accurate downstream OCR, validate model output on a small held-out set before bulk processing.
  • For archives mixing cards with other content, add negatives or apply a downstream non-card filter.
  • For cards in other languages, fine-tune with additional samples from those traditions before deployment.

Citation

BibTeX:

bibtex
@dataset{vanstrien_archival_index_cards_v3_2026,
  author    = {van Strien, Daniel},
  title     = {Archival Index Card Detection — mixed collections (v3)},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v3}
}

Dataset Card Authors

Daniel van Strien

Dataset Card Contact

Open an issue on the dataset repo or contact @davanstrien.