datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
institutional-books-hl-enriched-text
📚 Institutional Books: Harvard Library — Enriched Text
Institutional Books is a growing corpus of public domain books.
This release (IB-HL-ET) is a version of the text present in the Institutional Books: Harvard Library
(IB-HL) dataset
that has been further processed, filtered and optimized for computational access and model training.
This includes:
983K books, published largely in the 19th and 20th centuries
217B o200k_base tokens
7B sentences in 250 languages, grouped into… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-enriched-text.imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.archive-dolma3-mix-150b-enriched
archive-dolma3-mix-150b-enriched
ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B Dolma3 mix (not the pool).
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_mix_150B_enriched
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-mix-150b-enriched.geospatially_enriched_ndvi
Geospatially Enriched NDVI (16-Day Terra/MODIS)
This dataset transforms raw 16-day MODIS NDVI grids into a per-pixel time series enriched with hierarchical administrative boundaries. It covers every 0.1°×0.1° land pixel worldwide from 2000 onward and is partitioned for efficient bulk download and selective access.
Dataset Contents
Partitioned Parquet filesStored under:
ndvi/
├── year=YYYY/
│ ├── country=Netherlands/
│ │ └── data_0.parquet
│ └── country=India/
│ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/svenmeijboom/geospatially_enriched_ndvi.archive-dolma3-pool-150b-enriched
archive-dolma3-pool-150b-enriched
ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B pool sample.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_150B_enriched
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory including… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b-enriched.s2orc-cs-enriched
S2ORC CS Enriched
A Computer Science subset of the Semantic Scholar Open Research Corpus (S2ORC) enriched with LLM-generated structured metadata. Contains 1.1 million CS papers with extracted methods, models, datasets, metrics, compute estimates, and summaries.
Dataset Summary
Statistic
Value
Total papers
1,117,706
Total size
54.7 GB
Parquet files
1,118
Split
train
Dataset Structure
Base Columns
Content: parsed_title, abstract… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-cs-enriched.cifar100-enrichedThe CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).Biomed-Enriched
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
Dataset Authors
Rian Touchent, Nathan Godey & Eric de la ClergerieSorbonne Université, INRIA Paris
Overview
Biomed-Enriched is a PubMed-derived dataset created using a two-stage annotation process. Initially, Llama 3.1 70B Instruct annotated 400K paragraphs for document type, domain, and educational quality. These annotations were then… See the full description on the dataset page: https://huggingface.co/datasets/almanach/Biomed-Enriched.GridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.oxford-iiit-pet-vl-enriched
Visualize on Visual Layer
Oxford-IIIT-Pets-VL-Enriched
An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues!
With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues help to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.carbon-cpu-enriched-sequences-sampledspeech_commands_enriched_and_annotated
Dataset Summary
📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development.
🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways:
Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.BDD100K-enrichedipl-enriched-dataset
Dataset Summary
This dataset is an enriched version of the IPL Dataset 2008-2026.It starts from the original Kaggle data and adds analytics-driven, derived attributes to improve usefulness for machine learning and advanced data analysis workflows.
Modifications & Derived Attributes
The base data was extended with new engineered features created through extensive analytics.
The final enriched file adds 27 derived attributes:
match_phase - Phase bucket by over:… See the full description on the dataset page: https://huggingface.co/datasets/prasad-gade05/ipl-enriched-dataset.dcase23-task2-enriched
Dataset Card for the Enriched "DCASE 2023 Challenge Task 2 Dataset".
Dataset Summary
Data-centric AI principles have become increasingly important for real-world use cases. At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development.
This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps… See the full description on the dataset page: https://huggingface.co/datasets/renumics/dcase23-task2-enriched.coco-2014-vl-enriched
Visualize on Visual Layer
COCO-2014-VL-Enriched
An enriched version of the COCO 2014 dataset with label issues! The label issues help to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original image filename from the COCO dataset.
image: Image data in the form of PIL Image.
label_bbox: Bounding box annotations from the COCO dataset. Consists of bounding box coordinates, confidence scores, and labels… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/coco-2014-vl-enriched.inaturalist-enriched
Enriched iNaturalist dataset from 2026-03-27.
This dataset is based on philipp-zettl/inaturalist-s3-massive.
The data was enriched using the ./enrich.py script inside the repository.
It contains the following features
photo_id: The original ID of the photo inside the inaturalist dataset
observation_uuid: The observation's UUID
image: The actual image content
taxon_id: The ID of the taxonomy
species_name: The name of the species inside the image
taxonomic_rank: The type of taxonomic rank… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/inaturalist-enriched.doclaynet-pt-enriched-formulapl-nsa-enriched
Polish NSA Judgments (Enriched)
Polish Supreme Administrative Court judgments enriched with Gemini-extracted factual_state and legal_state fields.
Dataset Description
This dataset is an enriched version of JuDDGES/pl-nsa with additional fields extracted using Google Gemini 2.5 Pro.
New Fields
Core Extracted Fields
Field
Type
Description
factual_state
string
Objective narrative of facts (stan faktyczny) - the factual circumstances forming… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/pl-nsa-enriched.fineweb-enriched-classifiedlibrispeech_asr_enriched
Dataset Card for librispeech_asr_enriched
sms-spam-enriched
SMS Spam Enriched Dataset
An enriched version of the classic SMS Spam Collection Dataset from UC Irvine with additional engineered features and semantic embeddings.This dataset is designed for spam detection, feature engineering experiments, and model interpretability research.
Dataset Overview
Total samples: 5,171
Classes:
0: Ham (non-spam)
1: Spam
Enrichments Added
Alongside the raw SMS text (sms) and labels (label), we engineered multiple new… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/sms-spam-enriched.enriched_web_nlgWebNLG is a valuable resource and benchmark for the Natural Language Generation (NLG) community. However, as other NLG benchmarks, it only consists of a collection of parallel raw representations and their corresponding textual realizations. This work aimed to provide intermediate representations of the data for the development and evaluation of popular tasks in the NLG pipeline architecture (Reiter and Dale, 2000), such as Discourse Ordering, Lexicalization, Aggregation and Referring Expression Generation.processed-financial-news-XXL-enrichedchatalpaca-multiturn-enriched-2.1cifar10-enrichedThe CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images
per class. There are 50000 training images and 10000 test images.
This version if CIFAR-10 is enriched with several metadata such as embeddings, baseline results and label error scores.gloom-data-enriched-regeneratedspeech_commands_enrichedThis is a set of one-second .wav audio files, each containing a single spoken
English word or background noise. These words are from a small set of commands, and are spoken by a
variety of different speakers. This data set is designed to help train simple
machine learning models. This dataset is covered in more detail at
[https://arxiv.org/abs/1804.03209](https://arxiv.org/abs/1804.03209).
Version 0.01 of the data set (configuration `"v0.01"`) was released on August 3rd 2017 and contains
64,727 audio files.
In version 0.01 thirty different words were recoded: "Yes", "No", "Up", "Down", "Left",
"Right", "On", "Off", "Stop", "Go", "Zero", "One", "Two", "Three", "Four", "Five", "Six", "Seven", "Eight", "Nine",
"Bed", "Bird", "Cat", "Dog", "Happy", "House", "Marvin", "Sheila", "Tree", "Wow".
In version 0.02 more words were added: "Backward", "Forward", "Follow", "Learn", "Visual".
In both versions, ten of them are used as commands by convention: "Yes", "No", "Up", "Down", "Left",
"Right", "On", "Off", "Stop", "Go". Other words are considered to be auxiliary (in current implementation
it is marked by `True` value of `"is_unknown"` feature). Their function is to teach a model to distinguish core words
from unrecognized ones.
This version is not yet supported.
The `_silence_` class contains a set of longer audio clips that are either recordings or
a mathematical simulation of noise.chatalpaca-multiturn-enriched
