datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glami-1m-t2i-mteb
GLAMI-1M text-to-image retrieval
This MTEB-formatted derivative uses the complete 116,004-row official GLAMI-1M test split. Product names and descriptions are text queries and product images are the corpus. Repeated image IDs and exact repeated texts are deduplicated within each language, and qrels retain every observed text-image association.
The unchanged source archives are already hosted by the original authors in glami/glami-1m. GLAMI-1M-dataset--test-only.zip is pinned at… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-t2i-mteb.glami-1m
GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem.
Each sample contains an image, country code, name in corresponding language, description, target category and source of the label which can be of multiple types… See the full description on the dataset page: https://huggingface.co/datasets/glami/glami-1m.GLAMI-1M
This is fork of original dataset converted to dataset format.
GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem.
Each sample contains an image, country code, name in corresponding language… See the full description on the dataset page: https://huggingface.co/datasets/pySilver/GLAMI-1M.glami-1m-mteb
GLAMI-1M MTEB multimodal classification
This is an MTEB-ready derivative of the official
glami/glami-1m
release for multilingual image+text fashion classification. The source is
pinned at revision befda45d8d4e8b8082bb8a1912d1f9eb9483991c and remains
licensed under Apache-2.0.
Each example contains the official product image, name and description
joined as text, and the official category ID as label. The complete
116,004-row human-labeled test split is unchanged.
To keep… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-mteb.glam-extraction-benchmark
GLAM extraction benchmark
Structured extraction from cultural-heritage documents. The first configuration is
nls-index-cards: 98 manuscript catalogue cards from the National Library of Scotland.
Source and credits
Derived from NationalLibraryOfScotland/index-cards-eval,
revision 2a81070549d8493c2c538744a9dbbc1dc72cb146 (CC0). Images and checked outputs are preserved.
NLS cataloguers reviewed the model-drafted labels: 66 accepted as drafted, 32 corrected.
Drafting… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/glam-extraction-benchmark.bl-crop-tighten-v1
bl-crop-tighten-v1
Training data for crop tightening on the British Library Book Images collection: 7,565 ABBYY picture-block crops (train 6,050 / validation 757 / test 758) with instance boxes and segmentation masks. The splits are book-safe — no book appears in more than one split (4,484 books total).
The labels are weak labels, not human annotations: tiiuae/Falcon-Perception-0.6B ran open-vocabulary segmentation over 8,400 stratified crops (embellishments, plates, medium… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/bl-crop-tighten-v1.GLAMI-1M-remapped
This is fork of original dataset converted to dataset format with adjusted category names.
GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem.
Each sample contains an image, country code, name in… See the full description on the dataset page: https://huggingface.co/datasets/pySilver/GLAMI-1M-remapped.index-card-detection-v3
Dataset Card for Archival Index Card Detection — mixed collections
A training dataset for object detection of index cards in archival scans. Combines four publicly-released collections — NLS Advocates Library single-card pages, US Navy Nurse Corps multi-card biographical sheets, Boston Public Library catalog cards, and Duke Rubenstein manuscript catalog cards — into a single object-detection schema.
Dataset Details
Dataset Description
1,425 archival scans… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v3.index-card-detection-v5
Dataset Card for Archival Index Card Detection — v5 (ensemble-relabelled navy)
Refined version of small-models-for-glam/index-card-detection-v3. All NLS / BPL / Rubenstein rows are passed through unchanged. The 25 navy-nurse-corps rows have their bounding boxes re-labelled via a v3+v4 model ensemble plus human review, replacing the SAM3-only bootstrap from v3.
Dataset Details
Dataset Description
Same 1,425-row mixed-collection composition as v3. The… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v5.autotrain-zyermpvktmz6qr9uqy4xfu8xscu2zyermpvktmz6qr9uqy4xfu8xscu2index-card-blank-content
Index-card blank / content / divider classifier — dataset
Cropped single archival index cards labelled blank, content, or divider, for
training a tiny CPU pre-filter that skips blank/divider cards before expensive VLM metadata
extraction in card-catalogue digitisation pipelines.
Two collections: Boston Public Library (BPL) FRC shelf-list cards and National Library
of Scotland (NLS) Advocates Library cards. Styles differ, so evaluate per collection.
How it was made… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-blank-content.GLAMI-1M-test-onlyGLAMI-1M-convo-smallglami-1m
GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem.
Each sample contains an image, country code, name in corresponding language, description, target category and source of the label which can be of multiple types… See the full description on the dataset page: https://huggingface.co/datasets/Imogenhb/glami-1m.autotrain-ali-imagesautotrain-zyermpvktmz6qr9uqy4xfu8xscu2
