datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IPL-Cityscapes-Illuminants
IPL-CityscapesIlluminants-dataset
Illuminant modified dataset version of the famous autonomous driving semantic segmentation Cityscapes dataset.
Dataset generation
For each image, we generate a flat (constant) light spectrum and compute the pixel reflectances that obtain the RGM pixel values. Once we have the pixel reflectances, we generate different light spectrums of different dominant wavelengths (colors) and saturations and compute the new modified images. We apply… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/IPL-Cityscapes-Illuminants.cityscapesversion https://git-lfs.github.com/spec/v1
oid sha256:4bcf87ecfbbb8e07a01b21415a970c8b53a5283bf6872b657040d3f45c9241f7
size 31
IPL-Cityscapes-LuminanceContrasts
IPL-CityscapesLuminanceContrasts-dataset
Controled luminance and contrasts modified dataset version of the famous autonomous driving semantic segmentation Cityscapes dataset.
Dataset generation
For each original image, we convert it to ATD color space. Once in this space, we compute its mean luminance, achromatic contrast and chromatic contast. We modify each of its characteristics in turns from 0.5 to 1.5 of its original value. Then we return the image to RGB space. We… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/IPL-Cityscapes-LuminanceContrasts.TTA-Cityscapes-Ccityscape-adverse
Cityscape‑Adverse
A benchmark for evaluating semantic segmentation robustness under realistic adverse conditions.
Overview
Cityscape‑Adverse extends the original Cityscapes dataset by applying eight realistic environmental modifications—rainy, foggy, spring, autumn, snowy, sunny, night, and dawn—using diffusion‑based image editing. All transformations preserve the original 2048×1024 semantic labels, enabling direct evaluation of model robustness in… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cityscape-adverse.cityscapesWild-City
WildCity Dataset
WildCity is a real-world city-scale multimodal dataset for street-view reconstruction, simulation, and spatial intelligence. It is collected from autonomous-driving fleet logs across multiple U.S. cities and contains surround-view RGB images, LiDAR, calibration, ego and sensor poses, object annotations, semantic masks, and processed reconstruction assets.
This repository hosts the initial public release of WildCity. This version does not include the full raw… See the full description on the dataset page: https://huggingface.co/datasets/Neptune615/Wild-City.spaq-cityscape-indoor-scene
SPAQ Cityscape + Indoor Scene Export
Images are copied without modification from SPAQ.
Selection rule: Cityscape > 0 OR Indoor scene > 0 from Scene category labels.xlsx.
Corrupt or missing images are skipped and recorded in failed_files.jsonl.
City-Landscape-In-Sight
City Landscape In Sight — Window View Perception (Images & Models)
This dataset hosts the large binary artefacts for the paper "City landscape in sight: A crowdsourced framework for unlocking urban-scale window view perceptions from real estate imagery." It is the companion of the code repository on GitHub:
Paper (arXiv): https://arxiv.org/abs/2606.15198
Code & derived data (GitHub): https://github.com/Sijie-Yang/City-Landscape-In-Sight
Images & trained weights (this dataset):… See the full description on the dataset page: https://huggingface.co/datasets/sijiey/City-Landscape-In-Sight.china-city-osm-pbf
gene-osm
gene-osm generates city-level OpenStreetMap PBF files from a country- or
region-level PBF. Each output is identified by the six-digit Chinese
administrative code (adcode) from the Ministry of Civil Affairs' official
2022 county-and-above administrative division table.
Output filenames preserve official administrative names:
<adcode>_<official_name>.osm.pbf
140600_朔州市.osm.pbf
Extraction always uses an actual polygon boundary and Osmium's
complete_ways strategy. There is… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/china-city-osm-pbf.cityscapes_segmentationAeroIntent-city
AeroIntent: City Scene
This repository is one scene of AeroIntent, a synchronized multimodal benchmark for UAV intent recognition.
Review Status and Access
AeroIntent is currently under peer review. Access is manually gated during the review period, and the complete benchmark will be made public after the paper is accepted. Please do not redistribute the data or use it for publication before the public release.
Contents
15 UAV configurations from… See the full description on the dataset page: https://huggingface.co/datasets/AerointentBenchmark/AeroIntent-city.netryx-new-york-city-13km
Nyc-Core-Usethis 13km
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 40.713200, -74.002500
Radius: 13.0 km
Panoramas: 663,084
Index entries: 2,652,336
Descriptor model: MegaLoc
Descriptor dim: 1024 (PCA from 8448)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("nyc-core-usethis-13km", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-city-13km.CityLearnThe dataset consists of tuples of (observations, actions, rewards, dones) sampled by agents
interacting with the CityLearn 2022 Phase 1 environment (only first 5 buildings)historical-danish-handwriting
Dataset Description
Dataset Summary
The Historical Danish handwriting dataset is a Danish-language dataset containing more than 11.000 pages of transcribed and proofread handwritten text.
The dataset currently consists of the published minutes from a number of City and Parish Council meetings, all dated between 1841 and 1939.
Languages
All the text is in Danish. The BCP-47 code for Danish is da.
Dataset Structure
Data Instances
Each data… See the full description on the dataset page: https://huggingface.co/datasets/aarhus-city-archives/historical-danish-handwriting.us-property-tax-2026-state-county-city
US Property Tax Rates 2026 — State, County & City (13,241 jurisdictions)
Compiled by the Calcuris Data Team — calcuris.com · License: CC BY 4.0
The most granular open compilation of US residential property tax burdens we are aware of:
51 states/DC + 3,132 counties + 10,058 cities (Census places with population ≥ 2,500), each with
the effective property tax rate, median annual tax paid, median home value and median household income.
Files
File
Rows
Contents… See the full description on the dataset page: https://huggingface.co/datasets/tresor2k/us-property-tax-2026-state-county-city.modern-city-unreal-synthetic
Modern city junction: a synthetic detection capture rendered in Unreal Engine
300 rendered frames from a single Unreal Engine environment, carrying
6,807 labelled object instances. Every box and every mask here is read
out of the engine's own per-instance ID buffer at render time. No model produced these labels
and no one drew them by hand, so a label is wrong only where the scene description behind it
is wrong.
The capture
Engine… See the full description on the dataset page: https://huggingface.co/datasets/NameFrame/modern-city-unreal-synthetic.city_temperature_anomaliesgta5-cityscapes-labelingAnchor_imagesfloorplans-cityscapes
Dataset Summary
This is a curated collection of floorplan images sourced from across the internet. It is intended for research in architectural AI, layout generation, and urban scene understanding.
Data format: Image files with associated integer labels.
Sources: Publicly available images from various web sources (This dataset is one unified collections).
Purpose: Educational and research use.
Dataset Structure
The dataset follows the standard Hugging Face Image… See the full description on the dataset page: https://huggingface.co/datasets/wheres-my-python/floorplans-cityscapes.now-city-pop-clipCityCube-BenchCityscapesIQA_dataCityscapedead-city-unreal-synthetic
Dead city: a synthetic detection capture rendered in Unreal Engine
300 rendered frames from a single Unreal Engine environment, carrying
10,693 labelled object instances. Every box and every mask here is read
out of the engine's own per-instance ID buffer at render time. No model produced these labels
and no one drew them by hand, so a label is wrong only where the scene description behind it
is wrong.
The environment was built by George Shachnev, who sent us the map and agreed… See the full description on the dataset page: https://huggingface.co/datasets/NameFrame/dead-city-unreal-synthetic.cityscapes
CityScapes Dataset (HuggingFace Format)
Dataset Description
This is a processed version of the CityScapes dataset converted to HuggingFace datasets format. The dataset contains urban street scenes with corresponding semantic segmentation labels and depth maps, commonly used for autonomous driving research and computer vision tasks.
Dataset Summary
Total samples: 3,475 images
Training samples: 2,975 images
Validation samples: 500 images
Image resolution:… See the full description on the dataset page: https://huggingface.co/datasets/tanganke/cityscapes.Procedural-City-Multimodal-Dataset-V2
Synthetic Urban Multimodal Dataset - Volume 2 (v0.2)
This is the second major release of the Synthetic Urban Multimodal Dataset, part of an ongoing observation project on AI-generated data ecosystems.
Volume 2 introduces a completely new collection of urban environments to the series, providing 40,000 high-precision files (8,000 completely new unique scenes × 5 modalities) generated entirely via the "Constructive Furnace" (our custom Blender-Python pipeline).
What's New… See the full description on the dataset page: https://huggingface.co/datasets/jp-cypress/Procedural-City-Multimodal-Dataset-V2.cityscapesThis dataset is part of the CycleGAN datasets, originally hosted here: https://people.eecs.berkeley.edu/~taesung_park/CycleGAN/datasets/
Citation
@article{DBLP:journals/corr/ZhuPIE17,
author = {Jun{-}Yan Zhu and
Taesung Park and
Phillip Isola and
Alexei A. Efros},
title = {Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial
Networks},
journal = {CoRR},
volume = {abs/1703.10593},
year… See the full description on the dataset page: https://huggingface.co/datasets/huggan/cityscapes.
