CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artist /glami-1m-t2i-mteb GLAMI-1M text-to-image retrieval This MTEB-formatted derivative uses the complete 116,004-row official GLAMI-1M test split. Product names and descriptions are text queries and product images are the corpus. Repeated image IDs and exact repeated texts are deduplicated within each language, and qrels retain every observed text-image association. The unchanged source archives are already hosted by the original authors in glami/glami-1m. GLAMI-1M-dataset--test-only.zip is pinned at… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-t2i-mteb.image100K<n<1M0 likes430 downloads21d agoHugging Face02zidcenek /GLAMI-Entity-Matching-Dataset GLAMI Duplication Detection Product-duplicate detection over GLAMI e-commerce listings: ~1.3M product images plus multilingual titles, descriptions and attributes, with labelled groups of items that do or do not refer to the same physical product. Released under the Apache License 2.0 — see LICENSE. TODO: describe how the labels were produced. Structure Config Files Contents images images/shard-*.parquet itemId → image bytes, one row per product image… See the full description on the dataset page: https://huggingface.co/datasets/zidcenek/GLAMI-Entity-Matching-Dataset.image-classification1M<n<10M0 likes230 downloads1mo agoHugging Face03glami /glami-1m GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem. Each sample contains an image, country code, name in corresponding language, description, target category and source of the label which can be of multiple types… See the full description on the dataset page: https://huggingface.co/datasets/glami/glami-1m.image6 likes196 downloads4y agoHugging Face04pySilver /GLAMI-1M This is fork of original dataset converted to dataset format. GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem. Each sample contains an image, country code, name in corresponding language… See the full description on the dataset page: https://huggingface.co/datasets/pySilver/GLAMI-1M.image1M<n<10M1 likes102 downloads1y agoHugging Face05artist /glami-1m-mteb GLAMI-1M MTEB multimodal classification This is an MTEB-ready derivative of the official glami/glami-1m release for multilingual image+text fashion classification. The source is pinned at revision befda45d8d4e8b8082bb8a1912d1f9eb9483991c and remains licensed under Apache-2.0. Each example contains the official product image, name and description joined as text, and the official category ID as label. The complete 116,004-row human-labeled test split is unchanged. To keep… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-mteb.imageimage-classification100K<n<1M0 likes73 downloads23d agoHugging Face06small-models-for-glam /glam-extraction-benchmark GLAM extraction benchmark Structured extraction from cultural-heritage documents. The first configuration is nls-index-cards: 98 manuscript catalogue cards from the National Library of Scotland. Source and credits Derived from NationalLibraryOfScotland/index-cards-eval, revision 2a81070549d8493c2c538744a9dbbc1dc72cb146 (CC0). Images and checked outputs are preserved. NLS cataloguers reviewed the model-drafted labels: 66 accepted as drafted, 32 corrected. Drafting… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/glam-extraction-benchmark.imageimage-to-textn<1K0 likes58 downloads7d agoHugging Face07small-models-for-glam /bl-crop-tighten-v1 bl-crop-tighten-v1 Training data for crop tightening on the British Library Book Images collection: 7,565 ABBYY picture-block crops (train 6,050 / validation 757 / test 758) with instance boxes and segmentation masks. The splits are book-safe — no book appears in more than one split (4,484 books total). The labels are weak labels, not human annotations: tiiuae/Falcon-Perception-0.6B ran open-vocabulary segmentation over 8,400 stratified crops (embellishments, plates, medium… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/bl-crop-tighten-v1.imageimage-segmentation1K<n<10K1 likes50 downloads1mo agoHugging Face08pySilver /GLAMI-1M-remapped This is fork of original dataset converted to dataset format with adjusted category names. GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem. Each sample contains an image, country code, name in… See the full description on the dataset page: https://huggingface.co/datasets/pySilver/GLAMI-1M-remapped.image1M<n<10M0 likes42 downloads1y agoHugging Face09small-models-for-glam /index-card-detection-v3 Dataset Card for Archival Index Card Detection — mixed collections A training dataset for object detection of index cards in archival scans. Combines four publicly-released collections — NLS Advocates Library single-card pages, US Navy Nurse Corps multi-card biographical sheets, Boston Public Library catalog cards, and Duke Rubenstein manuscript catalog cards — into a single object-detection schema. Dataset Details Dataset Description 1,425 archival scans… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v3.imageobject-detection1K<n<10K0 likes34 downloads4mo agoHugging Face10small-models-for-glam /index-card-detection-v5 Dataset Card for Archival Index Card Detection — v5 (ensemble-relabelled navy) Refined version of small-models-for-glam/index-card-detection-v3. All NLS / BPL / Rubenstein rows are passed through unchanged. The 25 navy-nurse-corps rows have their bounding boxes re-labelled via a v3+v4 model ensemble plus human review, replacing the SAM3-only bootstrap from v3. Dataset Details Dataset Description Same 1,425-row mixed-collection composition as v3. The… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v5.imageobject-detection1K<n<10K0 likes26 downloads4mo agoHugging Face11small-models-for-glam /synthetic-aat-materials Synthetic AAT Materials Dataset Dataset Description This dataset contains 1000 synthetic examples of cultural heritage object descriptions paired with their materials as they would appear in the Getty Art & Architecture Thesaurus (AAT). The data is formatted for training conversational AI models, particularly Qwen3, to identify and extract materials from cultural heritage object descriptions. Dataset Structure Each example contains: messages: Conversation… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/synthetic-aat-materials.texttext-generation1K<n<10K0 likes25 downloads1y agoHugging Face12small-models-for-glam /synthetic-linkedart-production-destructiontext10K<n<100K0 likes23 downloads9mo agoHugging Face13small-models-for-glam /synthetic-parsed-names-yaml Dataset Card for Synthetic Parsed Names (YAML) This dataset contains approximately 500,000 synthetic examples of complex, unstructured historical names paired with their structured YAML equivalents. It is designed to fine-tune small open-source large language models (LLMs) to accurately parse cultural heritage name strings into isolated components (first names, last names, middle names, dates, titles, etc.) for de-duplication and structured data ingestion. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/synthetic-parsed-names-yaml.text100K<n<1M2 likes18 downloads5mo agoHugging Face14glampirelabs /autotrain-zyermpvktmz6qr9uqy4xfu8xscu2zyermpvktmz6qr9uqy4xfu8xscu2imagen<1K0 likes17 downloads2y agoHugging Face15small-models-for-glam /synthetic-linkedart-physical-characteristicstext10K<n<100K0 likes16 downloads9mo agoHugging Face16small-models-for-glam /index-card-blank-content Index-card blank / content / divider classifier — dataset Cropped single archival index cards labelled blank, content, or divider, for training a tiny CPU pre-filter that skips blank/divider cards before expensive VLM metadata extraction in card-catalogue digitisation pipelines. Two collections: Boston Public Library (BPL) FRC shelf-list cards and National Library of Scotland (NLS) Advocates Library cards. Styles differ, so evaluate per collection. How it was made… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-blank-content.imageimage-classificationn<1K0 likes15 downloads4mo agoHugging Face17vaclavkosar /GLAMI-1M-test-onlyimage0 likes14 downloads4y agoHugging Face18Falah /arabic_glamour_prompts Dataset Card for "arabic_glamour_prompts" More Information needed text10K<n<100K0 likes13 downloads3y agoHugging Face19small-models-for-glam /aat-real-world Real-World AAT Materials Dataset This dataset contains 189,523 real-world examples of cultural heritage object material descriptions paired with their corresponding Art & Architecture Thesaurus (AAT) material classifications. Dataset Description The dataset is designed for training models to extract material information from cultural heritage object descriptions. Each example consists of: Input: A real material description from cultural heritage collections Output:… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/aat-real-world.text100K<n<1M1 likes13 downloads1y agoHugging Face20glamourdrive /hardstyleaudion<1K0 likes13 downloads3mo agoHugging Face21pySilver /GLAMI-1M-convo-smallimage10K<n<100K0 likes12 downloads1y agoHugging Face22schneewolflabs /glamour276Mix of curated websites (html, css, js) generated by various models (mostly Claude Opus 4.8, GPT Codex 5.3, Gemini Flash) on random prompts generated by Claude using Glamour. textn<1K0 likes11 downloads3mo agoHugging Face23schneewolflabs /glamour-opusschneewolflabs/glamour169 and schneewolflabs/glamour276 filtered for Claude Opus generations only. textn<1K0 likes10 downloads3mo agoHugging Face24VVVVVVVVIC /glam-datasets GLAM Datasets Heterogeneous manipulation demonstrations for GLAM: Imitation from Heterogeneous Demonstrations using Grounded Latent-Action World Models (project page · code). Each task combines a small target set — action-labelled demonstrations on our dual-arm Kinova robot (Frank) in MuJoCo — with a larger, action-free auxiliary set collected by a floating UMI gripper, also in MuJoCo. Contents Archive Task Episodes Size stack_two.tar task 1 — stack two… See the full description on the dataset page: https://huggingface.co/datasets/VVVVVVVVIC/glam-datasets.1K<n<10K0 likes9 downloads2mo agoHugging Face25small-models-for-glam /synthetic-parsed-namestext100K<n<1M0 likes7 downloads1y agoHugging Face26schneewolflabs /glamour169Mix of curated websites (html, css, js) generated by various models (mostly Claude Opus 4.8, GPT Codex 5.3, Gemini Flash) on random prompts generated by Claude using Glamour. textn<1K0 likes7 downloads3mo agoHugging Face27Imogenhb /glami-1m GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem. Each sample contains an image, country code, name in corresponding language, description, target category and source of the label which can be of multiple types… See the full description on the dataset page: https://huggingface.co/datasets/Imogenhb/glami-1m.image10K<n<100K0 likes4 downloads5mo agoHugging Face28glampirelabs /autotrain-ali-imagesimagen<1K0 likes3 downloads2y agoHugging Face29glampirelabs /autotrain-zyermpvktmz6qr9uqy4xfu8xscu2imagen<1K0 likes2 downloads2y agoHugging Face30minimew /glamira-raw-data0 likes1 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.