datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wildfires-cems
Wildfires - CEMS
The dataset includes annotations for burned area delineation and land cover segmentation, with a focus on European soil.
The dataset is curated from various sources, including the Copernicus Emergency Management System (EMS) and Sentinel-2 feeds.
Repository: https://github.com/links-ads/burned-area-seg
Paper: https://paperswithcode.com/paper/robust-burned-area-delineation-through
Dataset Preparation
The dataset has been compressed into segmentented… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/wildfires-cems.illustrated_ads
19th Century US Newspaper Adverts — illustrated or text-only
549 advertisement images cut from digitised United States newspaper pages in the Library of Congress Chronicling America collection, each labelled text-only or illustrations. The adverts were located by Newspaper Navigator (LC Labs), which ran an object detection model over 16,358,041 Chronicling America pages to extract visual content; advertisements are one of its categories.
The sample is even by year, not by title:… See the full description on the dataset page: https://huggingface.co/datasets/biglam/illustrated_ads.Korean_Real_Estate_Ads_Dataset
Korean Real Estate Ads Dataset
This dataset contains high-resolution images of Korean real estate advertisements, including online listings, printed flyers, and billboard ads for properties such as apartments, houses, and commercial spaces. It is designed to support AI research in OCR, visual understanding, and property analysis.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Korean_Real_Estate_Ads_Dataset.ADS504-Image-Arrays
Description
Roughly 15k images evenly split between AI generated and human generated for the purpose of training models to detect AI generated content.
Images have been converted to numpy arrays and resized to 224x224x3 for alignment with CNN models available on Tensorflow
Sources
The following sources were used to collect the images
AI vs HumanOpen Images V7Laion - 400M
Dataset_Info
Features:
name: imagedtype: numpy arrays (224x224x3)
name:… See the full description on the dataset page: https://huggingface.co/datasets/tkbarb10/ADS504-Image-Arrays.lookupjet-adsb-optical-tracking-jets-airplanes-aviation-samplesnull
Lookup-Jet: Multimodal Aviation Tracking Dataset
Overview
The Lookup-Jet Dataset is an industry-grade, sensor-fusion dataset designed for advanced computer vision and machine learning tasks. It combines high-resolution visual tracking (bounding boxes and polygon segmentations) of aircraft with synchronized ADS-B kinematics, environmental conditions, and astronomical data.
This dataset is pre-formatted for immediate deployment across standard ML frameworks… See the full description on the dataset page: https://huggingface.co/datasets/Lookupjet/lookupjet-adsb-optical-tracking-jets-airplanes-aviation-samples.
