CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01institutional /institutional-books-hl-enriched-textgated 📚 Institutional Books: Harvard Library — Enriched Text Institutional Books is a growing corpus of public domain books. This release (IB-HL-ET) is a version of the text present in the Institutional Books: Harvard Library (IB-HL) dataset that has been further processed, filtered and optimized for computational access and model training. This includes: 983K books, published largely in the 19th and 20th centuries 217B o200k_base tokens 7B sentences in 250 languages, grouped into… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-enriched-text.100K<n<1M2 likes4.7k downloads1mo agoHugging Face02visual-layer /imagenet-1k-vl-enriched Visualize on Visual Layer Imagenet-1K-VL-Enriched An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues! With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering. The label issues helps to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.imageobject-detection1M<n<10M40 likes4.1k downloads2y agoHugging Face03HCAI-Lab-GT /archive-dolma3-mix-150b-enriched archive-dolma3-mix-150b-enriched ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B Dolma3 mix (not the pool). Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_mix_150B_enriched Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-mix-150b-enriched.0 likes3.7k downloads4mo agoHugging Face04svenmeijboom /geospatially_enriched_ndvi Geospatially Enriched NDVI (16-Day Terra/MODIS) This dataset transforms raw 16-day MODIS NDVI grids into a per-pixel time series enriched with hierarchical administrative boundaries. It covers every 0.1°×0.1° land pixel worldwide from 2000 onward and is partitioned for efficient bulk download and selective access. Dataset Contents Partitioned Parquet filesStored under: ndvi/ ├── year=YYYY/ │ ├── country=Netherlands/ │ │ └── data_0.parquet │ └── country=India/ │ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/svenmeijboom/geospatially_enriched_ndvi.text1B<n<10B0 likes3.6k downloads1y agoHugging Face05HCAI-Lab-GT /archive-dolma3-pool-150b-enriched archive-dolma3-pool-150b-enriched ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B pool sample. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_pool_150B_enriched Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory including… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b-enriched.0 likes3.1k downloads4mo agoHugging Face06AlgorithmicResearchGroup /s2orc-cs-enriched S2ORC CS Enriched A Computer Science subset of the Semantic Scholar Open Research Corpus (S2ORC) enriched with LLM-generated structured metadata. Contains 1.1 million CS papers with extracted methods, models, datasets, metrics, compute estimates, and summaries. Dataset Summary Statistic Value Total papers 1,117,706 Total size 54.7 GB Parquet files 1,118 Split train Dataset Structure Base Columns Content: parsed_title, abstract… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-cs-enriched.tabulartext-classification1M<n<10M3 likes2.1k downloads5mo agoHugging Face07renumics /cifar100-enrichedThe CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses. There are two labels per image - fine label (actual class) and coarse label (superclass).imageimage-classification10K<n<100K4 likes1.4k downloads3y agoHugging Face08almanach /Biomed-Enriched Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Dataset Authors Rian Touchent, Nathan Godey & Eric de la ClergerieSorbonne Université, INRIA Paris Overview Biomed-Enriched is a PubMed-derived dataset created using a two-stage annotation process. Initially, Llama 3.1 70B Instruct annotated 400K paragraphs for document type, domain, and educational quality. These annotations were then… See the full description on the dataset page: https://huggingface.co/datasets/almanach/Biomed-Enriched.tabulartext-classification100M<n<1B9 likes1.4k downloads2mo agoHugging Face09beta3 /GridCorpus_9M_Sudoku_Puzzles_Enriched ╔══════════════════════════════════════════════════════════════════════╗ ║ ║ ║ G R I D C O R P U S ║ ║ ║ ║ "004300209005009001070060043..." ║ ║ │ ║ ║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.tabularfeature-extraction1M<n<10M1 likes1.3k downloads7mo agoHugging Face10AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes721 downloads7d agoHugging Face11visual-layer /oxford-iiit-pet-vl-enriched Visualize on Visual Layer Oxford-IIIT-Pets-VL-Enriched An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues! With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering. The label issues help to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.imageimage-classification1K<n<10K9 likes693 downloads2y agoHugging Face12AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes649 downloads27d agoHugging Face13soerenray /speech_commands_enriched_and_annotated Dataset Summary 📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development. 🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways: Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.audio10K<n<100K2 likes566 downloads3y agoHugging Face14lance-format /BDD100K-enrichedimage10K<n<100K0 likes555 downloads6mo agoHugging Face15prasad-gade05 /ipl-enriched-dataset Dataset Summary This dataset is an enriched version of the IPL Dataset 2008-2026.It starts from the original Kaggle data and adds analytics-driven, derived attributes to improve usefulness for machine learning and advanced data analysis workflows. Modifications & Derived Attributes The base data was extended with new engineered features created through extensive analytics. The final enriched file adds 27 derived attributes: match_phase - Phase bucket by over:… See the full description on the dataset page: https://huggingface.co/datasets/prasad-gade05/ipl-enriched-dataset.tabulartabular-classification100K<n<1M0 likes548 downloads1mo agoHugging Face16renumics /dcase23-task2-enriched Dataset Card for the Enriched "DCASE 2023 Challenge Task 2 Dataset". Dataset Summary Data-centric AI principles have become increasingly important for real-world use cases. At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development. This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps… See the full description on the dataset page: https://huggingface.co/datasets/renumics/dcase23-task2-enriched.audio-classification1K<n<10K6 likes530 downloads3y agoHugging Face17visual-layer /coco-2014-vl-enriched Visualize on Visual Layer COCO-2014-VL-Enriched An enriched version of the COCO 2014 dataset with label issues! The label issues help to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: The original image filename from the COCO dataset. image: Image data in the form of PIL Image. label_bbox: Bounding box annotations from the COCO dataset. Consists of bounding box coordinates, confidence scores, and labels… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/coco-2014-vl-enriched.imageobject-detection100K<n<1M2 likes488 downloads2y agoHugging Face18philipp-zettl /inaturalist-enriched Enriched iNaturalist dataset from 2026-03-27. This dataset is based on philipp-zettl/inaturalist-s3-massive. The data was enriched using the ./enrich.py script inside the repository. It contains the following features photo_id: The original ID of the photo inside the inaturalist dataset observation_uuid: The observation's UUID image: The actual image content taxon_id: The ID of the taxonomy species_name: The name of the species inside the image taxonomic_rank: The type of taxonomic rank… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/inaturalist-enriched.text1M<n<10M1 likes389 downloads6mo agoHugging Face19docling-project /doclaynet-pt-enriched-formulaimage100K<n<1M2 likes344 downloads11mo agoHugging Face20JuDDGES /pl-nsa-enriched Polish NSA Judgments (Enriched) Polish Supreme Administrative Court judgments enriched with Gemini-extracted factual_state and legal_state fields. Dataset Description This dataset is an enriched version of JuDDGES/pl-nsa with additional fields extracted using Google Gemini 2.5 Pro. New Fields Core Extracted Fields Field Type Description factual_state string Objective narrative of facts (stan faktyczny) - the factual circumstances forming… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/pl-nsa-enriched.text1M<n<10M0 likes337 downloads8mo agoHugging Face21Giuliaoc /fineweb-enriched-classifiedtext100K<n<1M0 likes330 downloads27d agoHugging Face22notnotsamuel /librispeech_asr_enriched Dataset Card for librispeech_asr_enriched automatic-speech-recognition100K<n<1M0 likes326 downloads7mo agoHugging Face23GenAIDevTOProd /sms-spam-enriched SMS Spam Enriched Dataset An enriched version of the classic SMS Spam Collection Dataset from UC Irvine with additional engineered features and semantic embeddings.This dataset is designed for spam detection, feature engineering experiments, and model interpretability research. Dataset Overview Total samples: 5,171 Classes: 0: Ham (non-spam) 1: Spam Enrichments Added Alongside the raw SMS text (sms) and labels (label), we engineered multiple new… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/sms-spam-enriched.tabular1K<n<10K0 likes322 downloads1y agoHugging Face24ThiagoCF05 /enriched_web_nlgWebNLG is a valuable resource and benchmark for the Natural Language Generation (NLG) community. However, as other NLG benchmarks, it only consists of a collection of parallel raw representations and their corresponding textual realizations. This work aimed to provide intermediate representations of the data for the development and evaluation of popular tasks in the NLG pipeline architecture (Reiter and Dale, 2000), such as Discourse Ordering, Lexicalization, Aggregation and Referring Expression Generation.tabular-to-text1K<n<10K2 likes311 downloads3y agoHugging Face25mateiplescan /processed-financial-news-XXL-enriched0 likes301 downloads5mo agoHugging Face26BRlkl /chatalpaca-multiturn-enriched-2.1text1K<n<10K0 likes291 downloads3mo agoHugging Face27renumics /cifar10-enrichedThe CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. This version if CIFAR-10 is enriched with several metadata such as embeddings, baseline results and label error scores.image-classification10K<n<100K1 likes290 downloads3y agoHugging Face28aksh-n /gloom-data-enriched-regeneratedtabular1K<n<10K0 likes282 downloads8mo agoHugging Face29renumics /speech_commands_enrichedThis is a set of one-second .wav audio files, each containing a single spoken English word or background noise. These words are from a small set of commands, and are spoken by a variety of different speakers. This data set is designed to help train simple machine learning models. This dataset is covered in more detail at [https://arxiv.org/abs/1804.03209](https://arxiv.org/abs/1804.03209). Version 0.01 of the data set (configuration `"v0.01"`) was released on August 3rd 2017 and contains 64,727 audio files. In version 0.01 thirty different words were recoded: "Yes", "No", "Up", "Down", "Left", "Right", "On", "Off", "Stop", "Go", "Zero", "One", "Two", "Three", "Four", "Five", "Six", "Seven", "Eight", "Nine", "Bed", "Bird", "Cat", "Dog", "Happy", "House", "Marvin", "Sheila", "Tree", "Wow". In version 0.02 more words were added: "Backward", "Forward", "Follow", "Learn", "Visual". In both versions, ten of them are used as commands by convention: "Yes", "No", "Up", "Down", "Left", "Right", "On", "Off", "Stop", "Go". Other words are considered to be auxiliary (in current implementation it is marked by `True` value of `"is_unknown"` feature). Their function is to teach a model to distinguish core words from unrecognized ones. This version is not yet supported. The `_silence_` class contains a set of longer audio clips that are either recordings or a mathematical simulation of noise.audioaudio-classification10K<n<100K3 likes259 downloads3y agoHugging Face30BRlkl /chatalpaca-multiturn-enrichedtext1K<n<10K0 likes249 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.