datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gs-images-v319c_newspapers_images_altogs-images-v4british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.open-imagesDepth-Normal-Images-617Kdatacomp_large_vie_imagesThis repository contains images downloaded with img2dataset for minhnguyent546/datacomp_large_vie_filtered2.
nasa-mars-rover-images
NASA Mars Rover Image Catalog
Credit: NASA/JPL-Caltech/MSSS
Part of a dataset collection on Hugging Face.
Dataset description
The NASA Mars Rover Image Catalog contains metadata for every raw image captured by the Perseverance (Mars 2020) and Curiosity (MSL) rovers on the surface of Mars. Perseverance has been exploring Jezero Crater since February 2021, investigating an ancient river delta for signs of past microbial life and caching samples for future… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/nasa-mars-rover-images.hsc-jwst-images-high-snrWhole_Slide_Imageshsc-jwst-imagesLoRA-Merge-Imagesarxiv-chandra-ocr-2-include-images-first50-20260415
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Source paper IDs in input list: 27,584
Processed IDs recorded in state/processed_ids.txt: 50
Successes: 50
Partial successes: 0
Errors: 0
Next shard index: 10
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.PickaPic-images
Dataset Card for "PickaPic-images"
More Information needed
compound-kanji-imagesball-maze-images
Dataset Card for Sagar18/ball-maze-images
This dataset contains visual observations of a ball maze environment, derived from the original state-space dataset at notmahi/tutorial-ball-top-20.
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("Sagar18/ball-maze-images")
# Access images (they will be automatically decoded)
image = dataset[0]["image"]
state = dataset[0]["state"]
nationalmuseet-open-images
Nationalmuseet Open Images
This dataset is an independently harvested research dataset from Nationalmuseet Samlinger Online.
It contains metadata and optionally WebDataset image shards for Nationalmuseet asset records whose
rights.license is one of:
Public Domain
CC-BY
No known rights
Public Domain and CC-BY are the strict open-license subset. No known rights is kept as a
separate license bucket because Nationalmuseet says this label means that, to their best assessment,
the… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/nationalmuseet-open-images.reachy-pick-and-place-imagesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "reachy2",
"total_episodes": 25,
"total_frames": 6652,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/erl-hub/reachy-pick-and-place-images.pickapic_v1_no_images_training_sfw
Dataset Information
This is an SFW sanitized prompt only version of the PickAPic dataset, with 335,000 prompts and image URLs.
Citation Information
If you find this work useful, please cite:
@inproceedings{Kirstain2023PickaPicAO,
title={Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation},
author={Yuval Kirstain and Adam Polyak and Uriel Singer and Shahbuland Matiana and Joe Penna and Omer Levy},
year={2023}
}
LICENSE
MIT License… See the full description on the dataset page: https://huggingface.co/datasets/CarperAI/pickapic_v1_no_images_training_sfw.candels-labeled-imageskbl-open-images
Royal Danish Library Open COP Images
This dataset is an independently harvested research dataset from the Royal Danish Library
(Det Kgl. Bibliotek) COP/Digital Collections API.
The export intentionally excludes:
aerial photographs / Danmark set fra luften
newspapers, including Danske aviser 1666-1883
text-heavy COP editions such as books, letters, manuscripts, pamphlets, catalogues and printed matter
Included COP editions:
Billeder
Kort og Atlas as a separate… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/kbl-open-images.nationalmuseet-open-images-synthrealman_task_better_images_20260818_191421This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.rot6d_0",
"left_ee.rot6d_1",
"left_ee.rot6d_2",
"left_ee.rot6d_3",
"left_ee.rot6d_4"… See the full description on the dataset page: https://huggingface.co/datasets/kumarhans/realman_task_better_images_20260818_191421.PickaPic-images-6-3-2023
Dataset Card for "PickaPic-images-6-3-2023"
More Information needed
candels-unlabeled-and-labeled-imagesrobocasa_spatial_images_pretrain_atomic_CloseCabinetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "PandaOmron",
"total_episodes": 105,
"total_frames": 27754,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:105"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kimz1121/robocasa_spatial_images_pretrain_atomic_CloseCabinet.arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Baworsar1/british-library-book-images.arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416.
