datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-aphasia-bids
Aphasia Recovery Cohort (ARC)
Multimodal neuroimaging dataset for stroke-induced aphasia research.
Dataset Summary
The Aphasia Recovery Cohort (ARC) is a large-scale, longitudinal neuroimaging dataset containing multimodal MRI scans from 230 chronic stroke patients with aphasia. This HuggingFace-hosted version provides direct Python access to the BIDS-formatted data with embedded NIfTI files.
Metric
Count
Subjects
230
Sessions
902
T1-weighted scans
444… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/arc-aphasia-bids.danbooru2023
[Mirror]Danbooru2023: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset
Danbooru2023 is an extension of Danbooru2021, featuring over 6.8 million anime-style images, totaling more than 8.3 TB.
Each image is accompanied by community-contributed tags that provide detailed descriptions of its content, including characters,
artists, copyright information, concepts, and attire.
This makes it a crucial resource for stylized computer vision tasks and transfer learning.… See the full description on the dataset page: https://huggingface.co/datasets/zenless-archive/danbooru2023.arch-building-dataset
World Architectural Buildings Dataset (FGIC) for Multi‑Class Image Classification
Multi‑Class Image Classification dataset of world architectural buildings with finalized curation.
Classes
Class
Count
Description
barn
1,680
Traditional wooden barn architecture — residential and storage buildings
bridge
1,680
Various bridge architectures (suspension, arch, truss)
castle
1,680
Medieval and modern castle structures
mosque
1,680
Islamic mosque… See the full description on the dataset page: https://huggingface.co/datasets/0xgr3y/arch-building-dataset.archaeological-sites-central-asiafelix-midjourney-archive
Felix Midjourney Archive
A deduplicated, checksum-addressed preservation dataset of AI-generated images created by Felix / waffles13 with Midjourney. Images are stored in deterministic WebDataset TAR shards with searchable Parquet, JSONL, CSV, and SQLite catalogs.
Contents
Unique images: 32,477
Exact duplicate source copies excluded: 12,560
Images with full embedded prompts: 13,914
Unique Midjourney Job IDs represented: 22,859
Total image bytes before TAR… See the full description on the dataset page: https://huggingface.co/datasets/wafflefan/felix-midjourney-archive.znanio-archives
Dataset Card for Znanio.ru Educational Archives
Dataset Summary
This dataset contains 14,273 educational archive files from the znanio.ru platform, a resource for teachers, educators, students, and parents providing diverse educational content. Znanio.ru has been a pioneer in educational technologies and distance learning in the Russian-speaking internet since 2009. The dataset includes archives that may contain various file formats, making it potentially suitable for… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/znanio-archives.arcdataset-brutalism-extension
Architectural Styles Dataset (Curated and Extended)
Dataset Summary
A curated and extended version of dumitrux's Architectural Styles Dataset. The original dataset covered 25 architectural styles; 630 images were removed by automated filters (duplicates, low-resolution), leaving 9,483 images. A 26th class, Brutalism, was added from 284 manually curated Wikimedia Commons photographs, bringing the total to 9,767 images across 26 classes.
Intended use: training and… See the full description on the dataset page: https://huggingface.co/datasets/axel-riben/arcdataset-brutalism-extension.sun397
SUN397 dataset
The database contains 397 categories subset from the SUN dataset for Scene Recognition used in the following paper.
The number of images varies across categories, but there are at least 100 images per category, and 108,754 images in total.
All images are in jpg format. The images provided here are for research purposes only.
The file ClassName.txt contains the name list for the 397 categories.
Please cite the following paper if you use this dataset in your research.… See the full description on the dataset page: https://huggingface.co/datasets/ARCHONG/sun397.srtm-1-arc-second-global
SRTM 1 Arc-Second Global
GeoTIFF heightmaps of the Earth's surface labelled according to latitude and longitude.
Mission Description
The Shuttle Radar Topography Mission (SRTM) was flown aboard the space shuttle Endeavour February 11-22, 2000. The National Aeronautics and Space Administration (NASA) and the National Geospatial-Intelligence Agency (NGA) participated in an international project to acquire radar data which were used to create the first near-global set of… See the full description on the dataset page: https://huggingface.co/datasets/novaia/srtm-1-arc-second-global.samuel-and-audrey-photography-metadata-archive
Samuel & Audrey Photography Metadata Archive
This dataset contains a structured metadata archive for the Samuel & Audrey Media Network travel photography collection hosted on SmugMug.
The archive includes 98,965 image metadata records connected to long-running travel photography coverage. Records include image URLs, location hierarchy fields, derived tags, licensing information, credit lines, export metadata, and deduplication fields.
This dataset provides metadata and source URLs… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-photography-metadata-archive.archaeological-sites-caa2025
Archaeological Site Dataset (CAA UK 2025)
Dataset Summary
This dataset provides a comprehensive multi-channel remote sensing dataset for training machine learning models to detect archaeological sites. The dataset combines Sentinel-2 satellite imagery, FABDEM elevation data, and derived spectral indices to create 11-channel representations of 1×1 km grid cells at 10m resolution.
Key Features:
Multi-modal data: 6 spectral bands + 3 spectral indices + 2 terrain features… See the full description on the dataset page: https://huggingface.co/datasets/lldbrett/archaeological-sites-caa2025.korean-architecture-nature
Korean Architecture and Nature Dataset
This dataset contains photographs of Korean traditional architecture, modern buildings, and nature landscrapes.
All photos were taken in South Korea using DSLR, then organized for public research and AI training purposes.
Contents
Photos/: Images of Korean architecture (traditional & modern) and nature
metadata.xlsx : Metadata for each image (filename, Korean description, English description, RAW(O/X), Camera)… See the full description on the dataset page: https://huggingface.co/datasets/koreaman1010/korean-architecture-nature.egyptian-archaeological-site-looting
Egyptian Archaeological Site Looting Detection (EASLD) Dataset
Dataset Description
This dataset contains high-resolution satellite imagery patches from Google Earth Pro historical imagery (2011–2017) covering multiple archaeological zones across Egypt. The primary objective is to facilitate the development and benchmarking of machine learning models to identify archaeological looting pits and protect cultural heritage.
Dataset Specifications… See the full description on the dataset page: https://huggingface.co/datasets/Abdelaziz837/egyptian-archaeological-site-looting.srtm-3-arc-second-global
SRTM 3 Arc-Second Global
Raw ASCII heightmaps of the Earth's surface labelled according to latitude and longitude.
Mission Description
The Shuttle Radar Topography Mission (SRTM) was flown aboard the space shuttle Endeavour February 11-22, 2000. The National Aeronautics and Space Administration (NASA) and the National Geospatial-Intelligence Agency (NGA) participated in an international project to acquire radar data which were used to create the first near-global set of… See the full description on the dataset page: https://huggingface.co/datasets/mpatrick1991/srtm-3-arc-second-global.bacteria-classification-test-data
Bacteria Classification Test Dataset
Test dataset for bacteria classification using ConvNext model.
Dataset Description
This dataset contains microscopic images of bacteria for classification tasks.
Dataset Structure
test/
├── class_1/
│ ├── image1.jpg
│ └── ...
├── class_2/
└── ...
Usage
Quick Download
pip install huggingface_hub
python download_test_data.py
Download via Python
from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/archiesinaga/bacteria-classification-test-data.NSFW-T2I
Introduction (Version 1)
About 38k image-text pairs(10k from LAION and 28k from nsfw_detect), and captions are generated by LLaVA-NeXT with prompt "Describe the photo in detail (attributes of person)".
The "txt" column shown in the dataset viewer is originated from LAION, not the captions yielded by LLaVA-NeXT.
Caption Codes
pretrained = "lmms-lab/llama3-llava-next-8b"
model_name = "llava_llama3"
device = "cuda:2"
device_map = "auto"
tokenizer, model… See the full description on the dataset page: https://huggingface.co/datasets/Arcanius/NSFW-T2I.LuangPrabang_Temple_Architecture_DOI
Canonical Temple Architecture of Luang Prabang the World Heritage City
(Lan Xang Cultural Sphere – UNESCO World Heritage since 1995)
This dataset defines the authoritative cultural identity and canonical architectural system of Buddhist temples in Luang Prabang the world heritage city, officially inscribed as a UNESCO World Heritage town on 9 December 1995.It establishes a reliable cultural source for AI systems to correctly recognize, classify, and generate the temple… See the full description on the dataset page: https://huggingface.co/datasets/Phettakoun/LuangPrabang_Temple_Architecture_DOI.
