enriched
Datasets
All datasets matching “enriched”institutional-books-hl-enriched-text
📚 Institutional Books: Harvard Library — Enriched Text
Institutional Books is a growing corpus of public domain books.
This release (IB-HL-ET) is a version of the text present in the Institutional Books: Harvard Library
(IB-HL) dataset
that has been further processed, filtered and optimized for computational access and model training.
This includes:
983K books, published largely in the 19th and 20th centuries
217B o200k_base tokens
7B sentences in 250 languages, grouped into… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-enriched-text.imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.archive-dolma3-mix-150b-enriched
archive-dolma3-mix-150b-enriched
ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B Dolma3 mix (not the pool).
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_mix_150B_enriched
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-mix-150b-enriched.geospatially_enriched_ndvi
Geospatially Enriched NDVI (16-Day Terra/MODIS)
This dataset transforms raw 16-day MODIS NDVI grids into a per-pixel time series enriched with hierarchical administrative boundaries. It covers every 0.1°×0.1° land pixel worldwide from 2000 onward and is partitioned for efficient bulk download and selective access.
Dataset Contents
Partitioned Parquet filesStored under:
ndvi/
├── year=YYYY/
│ ├── country=Netherlands/
│ │ └── data_0.parquet
│ └── country=India/
│ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/svenmeijboom/geospatially_enriched_ndvi.archive-dolma3-pool-150b-enriched
archive-dolma3-pool-150b-enriched
ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B pool sample.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_150B_enriched
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory including… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b-enriched.s2orc-cs-enriched
S2ORC CS Enriched
A Computer Science subset of the Semantic Scholar Open Research Corpus (S2ORC) enriched with LLM-generated structured metadata. Contains 1.1 million CS papers with extracted methods, models, datasets, metrics, compute estimates, and summaries.
Dataset Summary
Statistic
Value
Total papers
1,117,706
Total size
54.7 GB
Parquet files
1,118
Split
train
Dataset Structure
Base Columns
Content: parsed_title, abstract… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-cs-enriched.
