CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01quejing /small-meishi0 likes2k downloads4mo agoHugging Face02quejing20 /smallmeishi0 likes1.7k downloads4mo agoHugging Face03quejing /small-meishi-pdf0 likes432 downloads4mo agoHugging Face04ellaampy /SmallMinesDS SmallMinesDS The gradual expansion of unregularized artisanal small-scale gold mining (ASGM) fuels environmental degradation and poses risk to miners and mining communities. To enforce sustainable mining, support reclamation initiatives and pave the way for understudying the impacts of mining, we present SmallMinesDS, a benchmark dataset for mapping artisanal small-scale gold mining from multi-sensor satellite images. The initial version of the dataset covers five districts in… See the full description on the dataset page: https://huggingface.co/datasets/ellaampy/SmallMinesDS.image1K<n<10K2 likes297 downloads1y agoHugging Face05agentlans /small-magpie Smaller Magpie A collection of smaller Magpie datasets compared to agentlans/magpie. For argilla/magpie-ultra-v0.1, only instructions rated as good or excellent were selected. output_quality corresponds to the original dataset’s score_difference, which is the gap between instruct model and base model responses as evaluated by a reward model. Please see the original dataset for details. Source Rows argilla/magpie-ultra-v0.1 43923 Mxode/Magpie-Pro-10K-GPT4o-mini10000 texttext-generation100K<n<1M0 likes177 downloads10mo agoHugging Face06arjhinety /small-mind-post-training-data small-mind-companion — post-training data Every corpus used to post-train a ~2B vision-language model (google/gemma-4-E2B-it) for long-horizon personalised companion dialogue, in the order it was used: LoRA SFT → LoRA DPO → on-policy distillation. Part of the OneBee Datasets collection. Contents Path Rows Schema sft/v0/{train,val}.jsonl 202 / 23 messages sft/v1/{train,val}.jsonl 2232 / 248 messages dpo/v0/{train,val}.jsonl 200 / 23 prompt, chosen… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/small-mind-post-training-data.text-generation1K<n<10K0 likes117 downloads12d agoHugging Face07if001 /smallm-corpus-textbook Dataset Card for "smallm-corpus-textbook" smollm-corpusのCosmopedia v2のセットが大きすぎたので、text-bookの一部を抽出 https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus text1M<n<10M0 likes113 downloads1y agoHugging Face08arjhinety /small-mind-pmb-v0 PMB v0 — Personalised Memory Benchmark An evaluation benchmark for long-horizon personalised memory in small language models. It asks whether a model can recall what a specific user told it across many sessions, and — the part most memory benchmarks skip — whether it can decline to answer when the memory does not contain the answer. Built for small-mind-companion, a study of how much of the long-horizon memory gap a ~2B multimodal model can close without scaling parameters. Part… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/small-mind-pmb-v0.textquestion-answeringn<1K0 likes72 downloads12d agoHugging Face09m0pper /Small-Multilingual-CorporaSmall Multilingual Pretraining Copora used in the ICML 2025 Paper: Banyan: Improved Representation Learning with Explicit Structure It contains the following languages: Afrikaans: af Amharic: am Arabic: ar English: en Spanish: es Hausa: ha Hindi: hi Indonesian: id Marathi: mr Telugu: te Each language contains roughly between 10-100 million tokens of pre-training data. English is a sourced from a subsample of human written English Wikpedia. Amharic is a subsample of a corpus from the hub… See the full description on the dataset page: https://huggingface.co/datasets/m0pper/Small-Multilingual-Corpora.text1M<n<10M0 likes64 downloads1y agoHugging Face10small-models-for-glam /glam-extraction-benchmark GLAM extraction benchmark Structured extraction from cultural-heritage documents. The first configuration is nls-index-cards: 98 manuscript catalogue cards from the National Library of Scotland. Source and credits Derived from NationalLibraryOfScotland/index-cards-eval, revision 2a81070549d8493c2c538744a9dbbc1dc72cb146 (CC0). Images and checked outputs are preserved. NLS cataloguers reviewed the model-drafted labels: 66 accepted as drafted, 32 corrected. Drafting… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/glam-extraction-benchmark.imageimage-to-textn<1K0 likes59 downloads11d agoHugging Face11if001 /smallm-corpus-story Dataset Card for "smallm-corpus-story" smollm-corpusのCosmopedia v2のセットが大きすぎたので、storyの一部を抽出 https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus text100K<n<1M1 likes46 downloads1y agoHugging Face12small-models-for-glam /synthetic-linkedart-production-destructiontext10K<n<100K0 likes45 downloads9mo agoHugging Face13small-models-for-glam /bl-crop-tighten-v1 bl-crop-tighten-v1 Training data for crop tightening on the British Library Book Images collection: 7,565 ABBYY picture-block crops (train 6,050 / validation 757 / test 758) with instance boxes and segmentation masks. The splits are book-safe — no book appears in more than one split (4,484 books total). The labels are weak labels, not human annotations: tiiuae/Falcon-Perception-0.6B ran open-vocabulary segmentation over 8,400 stratified crops (embellishments, plates, medium… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/bl-crop-tighten-v1.imageimage-segmentation1K<n<10K1 likes45 downloads1mo agoHugging Face14lidai926 /small-minute-1855b1 small-minute-1855b1 Synthetic sensors test data: 36 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/lidai926/small-minute-1855b1.tabularn<1K0 likes36 downloads15d agoHugging Face15arjhinety /small-mind-probe-sets small-mind-companion — probe sets Two small, unrun probe sets from small-mind-companion. Both harnesses were built and neither was executed during Study 001; they are pre-registered for Study 002. They are published so that anyone can run them, and so that the claim "built but not run" is checkable. Part of the OneBee Datasets collection. h22_judgment/ — abliteration and judgment quality (24 probes) H22: removing a model's general refusal direction increases… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/small-mind-probe-sets.text-generationn<1K0 likes36 downloads12d agoHugging Face16small-models-for-glam /synthetic-parsed-names-yaml Dataset Card for Synthetic Parsed Names (YAML) This dataset contains approximately 500,000 synthetic examples of complex, unstructured historical names paired with their structured YAML equivalents. It is designed to fine-tune small open-source large language models (LLMs) to accurately parse cultural heritage name strings into isolated components (first names, last names, middle names, dates, titles, etc.) for de-duplication and structured data ingestion. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/synthetic-parsed-names-yaml.text100K<n<1M2 likes35 downloads5mo agoHugging Face17small-models-for-glam /synthetic-linkedart-physical-characteristicstext10K<n<100K0 likes30 downloads9mo agoHugging Face18yoki123 /small_MIAtext1K<n<10K0 likes28 downloads2y agoHugging Face19small-models-for-glam /index-card-blank-content Index-card blank / content / divider classifier — dataset Cropped single archival index cards labelled blank, content, or divider, for training a tiny CPU pre-filter that skips blank/divider cards before expensive VLM metadata extraction in card-catalogue digitisation pipelines. Two collections: Boston Public Library (BPL) FRC shelf-list cards and National Library of Scotland (NLS) Advocates Library cards. Styles differ, so evaluate per collection. How it was made… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-blank-content.imageimage-classificationn<1K0 likes28 downloads4mo agoHugging Face20small-models-for-glam /aat-real-world Real-World AAT Materials Dataset This dataset contains 189,523 real-world examples of cultural heritage object material descriptions paired with their corresponding Art & Architecture Thesaurus (AAT) material classifications. Dataset Description The dataset is designed for training models to extract material information from cultural heritage object descriptions. Each example consists of: Input: A real material description from cultural heritage collections Output:… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/aat-real-world.text100K<n<1M1 likes27 downloads1y agoHugging Face21small-models-for-glam /synthetic-aat-materials Synthetic AAT Materials Dataset Dataset Description This dataset contains 1000 synthetic examples of cultural heritage object descriptions paired with their materials as they would appear in the Getty Art & Architecture Thesaurus (AAT). The data is formatted for training conversational AI models, particularly Qwen3, to identify and extract materials from cultural heritage object descriptions. Dataset Structure Each example contains: messages: Conversation… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/synthetic-aat-materials.texttext-generation1K<n<10K0 likes23 downloads1y agoHugging Face22KwabsHug /small-model-schema-gym Small Model Schema Gym Dataset Deterministically generated chat examples for first-pass JSON compliance on the project Dream brief and Safety plan contracts. Files train.jsonl: 500 training examples. validation.jsonl: 200 held-out examples. manifest.json: counts, seed, provenance, and overlap check. Each row contains: { "id": "stable example identifier", "messages": [ {"role": "system", "content": "..."}, {"role": "user", "content": "..."}… See the full description on the dataset page: https://huggingface.co/datasets/KwabsHug/small-model-schema-gym.texttext-generationn<1K0 likes23 downloads3mo agoHugging Face23if001 /smallm-corpus-story-jasmollm-corpusのCosmopedia v2のセットのstoryをqwen3 8Bで日本語に翻訳したもの https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus text1K<n<10K0 likes22 downloads1y agoHugging Face24clemsadand /small_mc4A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI.text-generationn<1K0 likes21 downloads2y agoHugging Face25scottgeng00 /olmo2_delta_gpt5_vs_smallmodelstext100K<n<1M1 likes20 downloads1y agoHugging Face26dhai1999 /small-model-roberta0 likes20 downloads7mo agoHugging Face27small-models-for-glam /index-card-detection-v3 Dataset Card for Archival Index Card Detection — mixed collections A training dataset for object detection of index cards in archival scans. Combines four publicly-released collections — NLS Advocates Library single-card pages, US Navy Nurse Corps multi-card biographical sheets, Boston Public Library catalog cards, and Duke Rubenstein manuscript catalog cards — into a single object-detection schema. Dataset Details Dataset Description 1,425 archival scans… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v3.imageobject-detection1K<n<10K0 likes19 downloads4mo agoHugging Face28small-models-for-glam /index-card-detection-v5 Dataset Card for Archival Index Card Detection — v5 (ensemble-relabelled navy) Refined version of small-models-for-glam/index-card-detection-v3. All NLS / BPL / Rubenstein rows are passed through unchanged. The 25 navy-nurse-corps rows have their bounding boxes re-labelled via a v3+v4 model ensemble plus human review, replacing the SAM3-only bootstrap from v3. Dataset Details Dataset Description Same 1,425-row mixed-collection composition as v3. The… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-detection-v5.imageobject-detection1K<n<10K0 likes18 downloads4mo agoHugging Face29sinequa /Small-msmarcotext1M<n<10M0 likes16 downloads2y agoHugging Face30minhaozhang /small_mbtitext100K<n<1M0 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.