datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TACOTACO is a benchmark for Python code generation, it includes 25443 problems and 1000 problems for train and test splits.taco_dataset
TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding
Dataset Versions
[1] Pre-released Version
Dataset links:
OneDrive: https://1drv.ms/f/s!Ap-t7dLl7BFUfmNkrHubnoo8LCs?e=1h0Xhe
BaiduNetDisk: https://pan.baidu.com/s/1gANrhzdUyvsUGXcDB4xMfQ?pwd=kg7j
Dataset Contents:
244 high-quality motions sequences spanning 137 <tool, action, object> triplets
206 High-resolution object models (10K~100K faces per object mesh)
Hand-object pose and mesh… See the full description on the dataset page: https://huggingface.co/datasets/mzhobro/taco_dataset.methaneset
MethaneSET: Unified Multi-Sensor Datasets for Satellite-Based Methane Plume Detection
Authors: Cesar Aybar, Julio Contreras, David Montero, Miguel D. Mahecha, Luis Gómez-Chova
Paper: Scientific Data (under review)
Methane is the second-largest driver of anthropogenic warming, and a disproportionate share of emissions comes from a small number of super-emitters detectable by satellite. MethaneSET provides analysis-ready datasets for methane plume detection spanning three… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/methaneset.cloudsen12
This dataset follows the TACO specification.
cloudsen12plus
Website: https://cloudsen12.github.io/
version: 1.1.2
The largest dataset of expert-labeled pixels for cloud and cloud shadow detection in Sentinel-2
CloudSEN12+ version 1.1.0 is a significant extension of the CloudSEN12 dataset, which doubles the number of
expert-reviewed labels, making it, by a large margin, the largest cloud detection dataset to
date for Sentinel-2. All labels from the previous version have… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/cloudsen12.taco_play_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "franka",
"total_episodes": 3242,
"total_frames": 213972,
"total_tasks": 403,
"total_videos": 6484,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:3242"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/taco_play_lerobot.SEN2NAIPv2
This dataset follows the TACO specification.
sen2naipv2
A large-scale dataset for Sentinel-2 Image Super-Resolution
The SEN2NAIPv2 dataset is an extension of SEN2NAIP,
containing 62,242 LR and HR image pairs, about 76% more images than the first version. The dataset files
are named sen2naipv2-unet-000{1..3}.part.taco. This dataset comprises synthetic RGBN NAIP bands at 2.5 and 10 meters,
degraded to corresponding Sentinel-2 images and a potential x4 factor. The degradation… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/SEN2NAIPv2.mc_tacoMC-TACO (Multiple Choice TemporAl COmmonsense) is a dataset of 13k question-answer
pairs that require temporal commonsense comprehension. A system receives a sentence
providing context information, a question designed to require temporal commonsense
knowledge, and multiple candidate answers. More than one candidate answer can be plausible.
The task is framed as binary classification: givent he context, the question,
and the candidate answer, the task is to determine whether the candidate
answer is plausible ("yes") or not ("no").taco-datasetsThis repo consists of the datasets used for the TaCo paper. There are four datasets:
Multilingual Alpaca-52K GPT-4 dataset
Multilingual Dolly-15K GPT-4 dataset
TaCo dataset
Multilingual Vicuna Benchmark dataset
We translated the first three datasets using Google Cloud Translation.
The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets.
If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/taco-datasets.TACO-verified
Introduction
This dataset contains verified solutions from the TACO dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed.
The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds.
Statistics in the training set
Dataset
# Problems
# Solutions
TACO
25443
1468722
TACO-verified
12898
1043251
Correct Ratio
50.69 %
71.03 %… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/TACO-verified.RMBench-taco-gemini
RMBench-taco-gemini
RMBench training episodes (9 tasks, 449 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator:
Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 25 frames) and labels every sampled frame under the
task-specific context (taco) induced for that task. Each tick carries the current subtask, the running textual memory and the
visual-memory operations (keyframe store / retrieval) that the online… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-taco-gemini.sen2venus
This dataset follows the TACO specification.
SEN2VENµS: A Dataset for the Training of Sentinel-2 Super-Resolution Algorithms
Description
SEN2VENµS is an open dataset used for super-resolution of Sentinel-2 images by leveraging simultaneous acquisitions with the VENµS satellite. The original dataset includes 10m and 20m cloud-free surface reflectance patches from Sentinel-2, with reference spatially-registered surface reflectance patches at 5m resolution… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/sen2venus.taco_playThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 3603,
"total_frames": 237798,
"total_tasks": 406,
"total_videos": 7206,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:3603"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/taco_play.TACO-hf
BEE-spoke-data/TACO-hf
Simple re-host of https://huggingface.co/datasets/BAAI/TACO but saved as hf dataset for ease of use.
Features:
DatasetDict({
"train": Dataset({
"features": [
"question",
"solutions",
"starter_code",
"input_output",
"difficulty",
"raw_tags",
"name",
"source",
"tags",
"skill_types",
"url",
"Expected Auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/TACO-hf.RoboDojo-taco-visual-gemini
RoboDojo-taco-visual-gemini
The visual-grounding variant of RoboDojo-taco-gemini: the same
RoboDojo long-horizon episodes (8 tasks, 800 episodes, 25 fps) with the same dense high-level labels (Gemini 3.7 Flash under the
task-specific context induced for each task), except that a target position leaves the label text and is drawn into the
low-level policy's keyframe slot.
In the source labels a target that words cannot identify is named by its image coordinates on the 0–1000… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RoboDojo-taco-visual-gemini.taco_play_testworldfloods
This dataset follows the TACO specification.
WorldFloods: A Global Dataset for Operational Flood Extent Segmentation
Database
WorldFloods is a public dataset containing pairs of Sentinel-2 multispectral images (Level-1C) and corresponding flood segmentation masks. As of its latest release, it comprises 509 flood events worldwide, requiring approximately 300 GB of storage if fully downloaded. The primary goal is to facilitate automatic flood mapping from… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/worldfloods.RMBench-taco-wodemo-gemini
RMBench-taco-wodemo-gemini
RMBench training episodes (9 tasks, 450 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator:
Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 25 frames) and labels every sampled frame under a
task-specific context (taco) for that task. Each tick carries the current subtask, the running textual memory and the
visual-memory operations (keyframe store / retrieval) that the online high-level… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-taco-wodemo-gemini.thinking_taco_play_lerobot_output_qwen3vltaco_dataset_resized
TACO Resized (512x376) — Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding
This is the resized version of the TACO dataset, with all allocentric videos and segmentation masks downscaled to a uniform 512x376 resolution (from native 4096x3000 / 2048x1500). Camera intrinsics are rescaled accordingly.
Why use this version?
The original TACO allocentric videos are 4096x3000, making training impractical without on-the-fly resizing. This version… See the full description on the dataset page: https://huggingface.co/datasets/mzhobro/taco_dataset_resized.taco-api-fixtures
TACO API Fixtures
A deterministic collection of 50 small TACO datasets whose payloads are all
Rumi files. The fixtures exercise TACO contracts, metadata hierarchies,
FOLDER and ZIP containers, ZIP partitioning, TACOCAT consolidation, generated
locations, and local or remote range reads.
They are API fixtures, not training data or a scientific benchmark. Spatial
and temporal metadata generated here is intentionally synthetic.
Matrix
The repository combines ten… See the full description on the dataset page: https://huggingface.co/datasets/asterisk-labs/taco-api-fixtures.terminal_bench_2_tasktrove_dq_taco_step15_30b_a3b_20260729_222705
Agent trace dataset
OpenCode/Harbor rollout traces from the MarinSkyRL run
rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260726-235656-574ba8, exported with
make_and_upload_trace_dataset --episodes last (the last episode of each trial — the rollouts
the policy was trained on).
Coverage
Built from the complete trial set on durable object storage, not from a local evidence bundle.
quantity
value
trial directories on object storage
21711
trials with a… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_taco_step15_30b_a3b_20260729_222705.TACO-Waste-RecognitionRMBench-taco-luna
RMBench-taco-luna
RMBench training episodes (9 tasks, 449 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator:
GPT-5.6 Luna (gpt-5.6-luna) reads the frames of each episode sampled every 25 frames as labelled images and labels every sampled frame under the
task-specific context (taco) induced for that task. Each tick carries the current subtask, the running textual memory and the
visual-memory operations (keyframe store / retrieval) that the… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-taco-luna.OXE_taco_play_embeddingsLanguage Table (LeRobot) — Embedding-Only Release
(DINOv3 + SigLIP2 image features; EmbeddingGemma task-text features)
This repository packages a re-encoded variant of IPEC-COMMUNITY/taco_play_lerobot where raw videos are replaced by fixed-length image embeddings, and task strings are augmented with text embeddings. All indices, splits, and semantics remain consistent with the source dataset while storage and I/O are substantially lighter. To make the dataset practical to upload/download and… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/OXE_taco_play_embeddings.TACO-Benchmark
TACO-Benchmark
TACO (Text-to-SQL with Ambiguous and Cross-database Open-domain queries) is a benchmark for real-world data-lake Text-to-SQL.
📢 News (2026): TACO has been accepted to VLDB 2026! 🎉📄 Paper: arXiv:2606.14201
GitHub (code & evaluation): Akanezora0/TACO-Benchmark
Google Drive mirror: TACO-Benchmark.zip
Overview
Unlike Spider or BIRD — where the target database is known and schemas are clean — TACO evaluates systems on three challenges common in… See the full description on the dataset page: https://huggingface.co/datasets/Akanezora/TACO-Benchmark.taco
View on Pictograph · Pictograph Research · Creative Commons Attribution 4.0
About
TACO is a computer-vision dataset curated and annotated on Pictograph. The most common detected objects are grass, sidewalk, bush, bottle, frisbee, ruins. On Pictograph you can browse every annotated image, fork it into your own workspace in one click, export it in a dozen formats, or train a model on it directly.
At a glance
Metric
Value
Images
1,500
Annotations
4… See the full description on the dataset page: https://huggingface.co/datasets/pictograph/taco.tacos-captioningtaco2-checkpointspi05-libero-taco-evalRoboDojo-taco-gemini
RoboDojo-taco-gemini
RoboDojo long-horizon episodes (8 tasks, 800 episodes, 25 fps) with dense high-level labels produced by the TACOR offline annotator:
Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 25 frames) and labels every sampled frame under the
task-specific context (taco) induced for that task, in its hybrid form: labels name the target object's image coordinates only
where words cannot identify it (a random instance of a class that… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RoboDojo-taco-gemini.
