CoolFace
20 results

small-models

small-models-for-glam /glam-extraction-benchmark GLAM extraction benchmark Structured extraction from cultural-heritage documents. The first configuration is nls-index-cards: 98 manuscript catalogue cards from the National Library of Scotland. Source and credits Derived from NationalLibraryOfScotland/index-cards-eval, revision 2a81070549d8493c2c538744a9dbbc1dc72cb146 (CC0). Images and checked outputs are preserved. NLS cataloguers reviewed the model-drafted labels: 66 accepted as drafted, 32 corrected. Drafting… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/glam-extraction-benchmark.imageimage-to-textn<1K0 likes58 downloads8d agoHugging Facesmall-models-for-glam /bl-crop-tighten-v1 bl-crop-tighten-v1 Training data for crop tightening on the British Library Book Images collection: 7,565 ABBYY picture-block crops (train 6,050 / validation 757 / test 758) with instance boxes and segmentation masks. The splits are book-safe — no book appears in more than one split (4,484 books total). The labels are weak labels, not human annotations: tiiuae/Falcon-Perception-0.6B ran open-vocabulary segmentation over 8,400 stratified crops (embellishments, plates, medium… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/bl-crop-tighten-v1.imageimage-segmentation1K<n<10K1 likes53 downloads1mo agoHugging Facesmall-models-for-glam /synthetic-linkedart-production-destructiontext10K<n<100K0 likes40 downloads9mo agoHugging Facesmall-models-for-glam /index-card-detection-v3 Dataset Card for Archival Index Card Detection — mixed collections A training dataset for object detection of index cards in archival scans. Combines four publicly-released collections — NLS Advocates Library single-card pages, US Navy Nurse Corps multi-card biographical sheets, Boston Public Library catalog cards, and Duke Rubenstein manuscript catalog cards — into a single object-detection schema. Dataset Details Dataset Description 1,425 archival scans… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v3.imageobject-detection1K<n<10K0 likes35 downloads4mo agoHugging Facesmall-models-for-glam /synthetic-parsed-names-yaml Dataset Card for Synthetic Parsed Names (YAML) This dataset contains approximately 500,000 synthetic examples of complex, unstructured historical names paired with their structured YAML equivalents. It is designed to fine-tune small open-source large language models (LLMs) to accurately parse cultural heritage name strings into isolated components (first names, last names, middle names, dates, titles, etc.) for de-duplication and structured data ingestion. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/synthetic-parsed-names-yaml.text100K<n<1M2 likes30 downloads5mo agoHugging Facesmall-models-for-glam /index-card-detection-v5 Dataset Card for Archival Index Card Detection — v5 (ensemble-relabelled navy) Refined version of small-models-for-glam/index-card-detection-v3. All NLS / BPL / Rubenstein rows are passed through unchanged. The 25 navy-nurse-corps rows have their bounding boxes re-labelled via a v3+v4 model ensemble plus human review, replacing the SAM3-only bootstrap from v3. Dataset Details Dataset Description Same 1,425-row mixed-collection composition as v3. The… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v5.imageobject-detection1K<n<10K0 likes27 downloads4mo agoHugging Face