datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEN2NAIPv2
This dataset follows the TACO specification.
sen2naipv2
A large-scale dataset for Sentinel-2 Image Super-Resolution
The SEN2NAIPv2 dataset is an extension of SEN2NAIP,
containing 62,242 LR and HR image pairs, about 76% more images than the first version. The dataset files
are named sen2naipv2-unet-000{1..3}.part.taco. This dataset comprises synthetic RGBN NAIP bands at 2.5 and 10 meters,
degraded to corresponding Sentinel-2 images and a potential x4 factor. The degradation… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/SEN2NAIPv2.taco_play_testthinking_taco_play_lerobot_output_qwen3vlTACO-Waste-Recognitiontaco
View on Pictograph · Pictograph Research · Creative Commons Attribution 4.0
About
TACO is a computer-vision dataset curated and annotated on Pictograph. The most common detected objects are grass, sidewalk, bush, bottle, frisbee, ruins. On Pictograph you can browse every annotated image, fork it into your own workspace in one click, export it in a dozen formats, or train a model on it directly.
At a glance
Metric
Value
Images
1,500
Annotations
4… See the full description on the dataset page: https://huggingface.co/datasets/pictograph/taco.TACOThis is the dataset from http://tacodataset.org , cleaned up and repackaged.
The original is the COCO_format.zip, the Torch_format.zip is the same dataset transformed into torch format.
Both can be plugged directly into YOLO for fine tuning.
tortilla_demogeobench-vlm-ref-seg-rgbi-taco
GEOBench-VLM: Referring expression segmentation (four-band RGB+NIR)
331 referring expressions over 50 images, each resolving to its own binary mask (3-9 per image).
50 samples · splits: test 50 · tasks: referring-segmentation
Packaged as TACO v3.
Full description
Sample. The sample is the image, and the masks are a MASK_SET variable leaf, since N is the number of expressions rather than a time axis. The five prompt paraphrases per expression are… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/geobench-vlm-ref-seg-rgbi-taco.DeepExtremeCubes-video
This dataset follows the TACO specification.
Important Note: This dataset lives here now: https://huggingface.co/datasets/isp-uv-es/DeepExtremeCubes-video
DeepExtremeCubes-video: Sentinel-2 Minicubes in Video Format for Compound-Extreme Analysis
📝 Description
📦 Dataset
DeepExtremeCubes-video is a storage-efficient, analysis-ready re-packaging of the original DeepExtremeCubes collection.
All 42 k Sentinel-2 minicubes (2.56 km × 2.56 km, 2016-2022, 7… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/DeepExtremeCubes-video.magicbathynet-s2-pixel-class-taco
MagicBathyNet s2 (pixel class)
533 patches of 18x18 Sentinel-2 patches at a measured 9.97-9.98 m, each covering 180x180 m, over Agia Napa, Cyprus (35) and Puck Lagoon, Poland (498).
533 samples · splits: test 107 · train 364 · validation 62 · tasks: semantic-segmentation
Packaged as TACO v3.
Full description
Splits. The release's own list for this sensor and task: 426 train, 107 test, disjoint.
Target. A seabed class mask over five classes: poseidonia… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/magicbathynet-s2-pixel-class-taco.rsvqa-lr-taco
RSVQA-LR (2k evaluation subset)
2000 closed-form question/answer pairs over 100 Sentinel-2 tiles of 256x256 px at 10 m: presence, counting, comparison and rural/urban questions from the 2020 benchmark that defined remote-sensing VQA.
100 samples · splits: validation 100 · tasks: visual-question-answering
Packaged as TACO v3.
Full description
Scope. A 2,000-row evaluation subset of the release's validation split, not the ~77k-pair corpus. The full… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/rsvqa-lr-taco.satin-brazilian-coffee-scenes-taco
Brazilian Coffee Scenes (SATIN mirror)
Coffee against not-coffee in 2,876 SPOT tiles of 64x64 over four counties of Minas Gerais. A single-crop discrimination task: the negative class is every other land cover rather than another crop, and coffee is a woody perennial whose spectral signature sits close to natural arboreal vegetation. The three channels are green, red and near-infrared, not RGB.
2,876 samples · splits: test 572 · train 2,016 · validation… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/satin-brazilian-coffee-scenes-taco.idtrees-taco
IDTReeS 2020 individual tree crowns
Individual tree crown delineation and species identification in NEON airborne imagery: 85 plots of 20x20 m at two sites, each with a 0.1 m RGB orthophoto and a 1 m canopy height model.
85 samples · splits: train 68 · validation 17 · tasks: instance-segmentation, object-detection
Packaged as TACO v3.
Full description
Annotations. 1312 hand-delineated crown polygons, of which 1213 carry one of 33 field-identified… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/idtrees-taco.hyperspectral-sim-s2-waters-taco
Hyperspectral simulations for Sentinel-2 waters
266,471 simulated water-leaving remote sensing reflectance spectra with the inherent optical properties that generated them. Forward-modelled with the Bi et al. (2023) bio-geo-optical model through the WaterQuality framework (Koenig et al. 2023) and convolved to eight Sentinel-2 MSI bands (B1-B7, B8A).
266,471 samples · splits: test 26,647 · train 213,177 · validation 26,647 · tasks: regression
Packaged as… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/hyperspectral-sim-s2-waters-taco.TACO-Reformatted-Full
Dataset Card for "TACO-Reformatted-Full"
More Information needed
deepfake-detection-2026-images
TacoGido/deepfake-detection-2026-images
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('TacoGido/deepfake-detection-2026-images')
nwpu-vhr10-taco
NWPU VHR-10
800 very-high-resolution optical images (Google Earth and ISPRS Vaihingen) for geospatial object detection over 10 classes.
800 samples · splits: test 110 · train 575 · validation 115 · tasks: object-detection, instance-segmentation
Packaged as TACO v3.
Full description
Contents. 650 annotated images carrying 3921 objects with both boxes and instance polygons, plus 150 background images with no objects, kept as the explicit hard negatives… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/nwpu-vhr10-taco.satin-aid-multilabel-taco
AID MultiLabel (SATIN mirror)
3,000 aerial scenes at 600x600 from AID, relabelled with the 17 object and cover categories present in each -- airplane, cars, dock, mobile home, ship, tanks -- rather than with the one scene class AID itself assigns. The legend is the same 17 as dlrsd's, over different pixels and at 600 px instead of 256, so the two are a matched pair for asking whether a multi-label head generalises across resolution and source.
3,000… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/satin-aid-multilabel-taco.TACO_YOLO_Demo_assets
Assets for the TACO_YOLO_Demo
Demo: https://huggingface.co/spaces/fabiocigaina/TACO_YOLO_Demo
C-17audreyaudCatImagesTACO_Test_Reformatted
Dataset Card for "TACO_Test_Reformatted"
More Information needed
taco-robot-datasettacolispixelart
