CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BAAI /TACOTACO is a benchmark for Python code generation, it includes 25443 problems and 1000 problems for train and test splits.text-generation10K<n<100K144 likes16k downloads2y agoHugging Face02mzhobro /taco_dataset TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding Dataset Versions [1] Pre-released Version Dataset links: OneDrive: https://1drv.ms/f/s!Ap-t7dLl7BFUfmNkrHubnoo8LCs?e=1h0Xhe BaiduNetDisk: https://pan.baidu.com/s/1gANrhzdUyvsUGXcDB4xMfQ?pwd=kg7j Dataset Contents: 244 high-quality motions sequences spanning 137 <tool, action, object> triplets 206 High-resolution object models (10K~100K faces per object mesh) Hand-object pose and mesh… See the full description on the dataset page: https://huggingface.co/datasets/mzhobro/taco_dataset.video0 likes10k downloads6mo agoHugging Face03tacofoundation /methaneset MethaneSET: Unified Multi-Sensor Datasets for Satellite-Based Methane Plume Detection Authors: Cesar Aybar, Julio Contreras, David Montero, Miguel D. Mahecha, Luis Gómez-Chova Paper: Scientific Data (under review) Methane is the second-largest driver of anthropogenic warming, and a disproportionate share of emissions comes from a small number of super-emitters detectable by satellite. MethaneSET provides analysis-ready datasets for methane plume detection spanning three… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/methaneset.geospatialimage-segmentation100K<n<1M4 likes7.2k downloads2d agoHugging Face04tacofoundation /cloudsen12 This dataset follows the TACO specification. cloudsen12plus Website: https://cloudsen12.github.io/ version: 1.1.2 The largest dataset of expert-labeled pixels for cloud and cloud shadow detection in Sentinel-2 CloudSEN12+ version 1.1.0 is a significant extension of the CloudSEN12 dataset, which doubles the number of expert-reviewed labels, making it, by a large margin, the largest cloud detection dataset to date for Sentinel-2. All labels from the previous version have… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/cloudsen12.geospatial2 likes6.4k downloads2y agoHugging Face05IPEC-COMMUNITY /taco_play_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "franka", "total_episodes": 3242, "total_frames": 213972, "total_tasks": 403, "total_videos": 6484, "total_chunks": 4, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:3242" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/taco_play_lerobot.tabularrobotics100K<n<1M0 likes3.4k downloads2y agoHugging Face06tacofoundation /SEN2NAIPv2 This dataset follows the TACO specification. sen2naipv2 A large-scale dataset for Sentinel-2 Image Super-Resolution The SEN2NAIPv2 dataset is an extension of SEN2NAIP, containing 62,242 LR and HR image pairs, about 76% more images than the first version. The dataset files are named sen2naipv2-unet-000{1..3}.part.taco. This dataset comprises synthetic RGBN NAIP bands at 2.5 and 10 meters, degraded to corresponding Sentinel-2 images and a potential x4 factor. The degradation… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/SEN2NAIPv2.imagen<1K5 likes2.9k downloads2y agoHugging Face07CogComp /mc_tacoMC-TACO (Multiple Choice TemporAl COmmonsense) is a dataset of 13k question-answer pairs that require temporal commonsense comprehension. A system receives a sentence providing context information, a question designed to require temporal commonsense knowledge, and multiple candidate answers. More than one candidate answer can be plausible. The task is framed as binary classification: givent he context, the question, and the candidate answer, the task is to determine whether the candidate answer is plausible ("yes") or not ("no").question-answering10K<n<100K3 likes2.6k downloads3y agoHugging Face08saillab /taco-datasetsThis repo consists of the datasets used for the TaCo paper. There are four datasets: Multilingual Alpaca-52K GPT-4 dataset Multilingual Dolly-15K GPT-4 dataset TaCo dataset Multilingual Vicuna Benchmark dataset We translated the first three datasets using Google Cloud Translation. The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets. If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/taco-datasets.text1M<n<10M17 likes1.5k downloads3y agoHugging Face09likaixin /TACO-verified Introduction This dataset contains verified solutions from the TACO dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed. The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds. Statistics in the training set Dataset # Problems # Solutions TACO 25443 1468722 TACO-verified 12898 1043251 Correct Ratio 50.69 % 71.03 %… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/TACO-verified.textquestion-answering10K<n<100K20 likes1.5k downloads1y agoHugging Face10Myungkyu /RMBench-taco-gemini RMBench-taco-gemini RMBench training episodes (9 tasks, 449 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator: Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 25 frames) and labels every sampled frame under the task-specific context (taco) induced for that task. Each tick carries the current subtask, the running textual memory and the visual-memory operations (keyframe store / retrieval) that the online… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-taco-gemini.videorobotics1K<n<10K0 likes1.4k downloads17d agoHugging Face11tacofoundation /sen2venus This dataset follows the TACO specification. SEN2VENµS: A Dataset for the Training of Sentinel-2 Super-Resolution Algorithms Description SEN2VENµS is an open dataset used for super-resolution of Sentinel-2 images by leveraging simultaneous acquisitions with the VENµS satellite. The original dataset includes 10m and 20m cloud-free surface reflectance patches from Sentinel-2, with reference spatially-registered surface reflectance patches at 5m resolution… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/sen2venus.1 likes955 downloads2y agoHugging Face12lerobot /taco_playThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 3603, "total_frames": 237798, "total_tasks": 406, "total_videos": 7206, "total_chunks": 4, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:3603" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/taco_play.tabularrobotics100K<n<1M3 likes827 downloads1y agoHugging Face13BEE-spoke-data /TACO-hf BEE-spoke-data/TACO-hf Simple re-host of https://huggingface.co/datasets/BAAI/TACO but saved as hf dataset for ease of use. Features: DatasetDict({ "train": Dataset({ "features": [ "question", "solutions", "starter_code", "input_output", "difficulty", "raw_tags", "name", "source", "tags", "skill_types", "url", "Expected Auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/TACO-hf.texttext-generation10K<n<100K1 likes612 downloads9mo agoHugging Face14Myungkyu /RoboDojo-taco-visual-gemini RoboDojo-taco-visual-gemini The visual-grounding variant of RoboDojo-taco-gemini: the same RoboDojo long-horizon episodes (8 tasks, 800 episodes, 25 fps) with the same dense high-level labels (Gemini 3.7 Flash under the task-specific context induced for each task), except that a target position leaves the label text and is drawn into the low-level policy's keyframe slot. In the source labels a target that words cannot identify is named by its image coordinates on the 0–1000… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RoboDojo-taco-visual-gemini.videorobotics1K<n<10K0 likes506 downloads11d agoHugging Face15IrvingF7 /taco_play_testimage100K<n<1M0 likes444 downloads1y agoHugging Face16tacofoundation /worldfloods This dataset follows the TACO specification. WorldFloods: A Global Dataset for Operational Flood Extent Segmentation Database WorldFloods is a public dataset containing pairs of Sentinel-2 multispectral images (Level-1C) and corresponding flood segmentation masks. As of its latest release, it comprises 509 flood events worldwide, requiring approximately 300 GB of storage if fully downloaded. The primary goal is to facilitate automatic flood mapping from… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/worldfloods.1 likes439 downloads1y agoHugging Face17Myungkyu /RMBench-taco-wodemo-gemini RMBench-taco-wodemo-gemini RMBench training episodes (9 tasks, 450 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator: Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 25 frames) and labels every sampled frame under a task-specific context (taco) for that task. Each tick carries the current subtask, the running textual memory and the visual-memory operations (keyframe store / retrieval) that the online high-level… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-taco-wodemo-gemini.videorobotics1K<n<10K0 likes416 downloads9d agoHugging Face18ShareLab-SII /thinking_taco_play_lerobot_output_qwen3vlimage100K<n<1M0 likes382 downloads6mo agoHugging Face19mzhobro /taco_dataset_resized TACO Resized (512x376) — Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding This is the resized version of the TACO dataset, with all allocentric videos and segmentation masks downscaled to a uniform 512x376 resolution (from native 4096x3000 / 2048x1500). Camera intrinsics are rescaled accordingly. Why use this version? The original TACO allocentric videos are 4096x3000, making training impractical without on-the-fly resizing. This version… See the full description on the dataset page: https://huggingface.co/datasets/mzhobro/taco_dataset_resized.video0 likes370 downloads6mo agoHugging Face20asterisk-labs /taco-api-fixtures TACO API Fixtures A deterministic collection of 50 small TACO datasets whose payloads are all Rumi files. The fixtures exercise TACO contracts, metadata hierarchies, FOLDER and ZIP containers, ZIP partitioning, TACOCAT consolidation, generated locations, and local or remote range reads. They are API fixtures, not training data or a scientific benchmark. Spatial and temporal metadata generated here is intentionally synthetic. Matrix The repository combines ten… See the full description on the dataset page: https://huggingface.co/datasets/asterisk-labs/taco-api-fixtures.geospatialn<1K0 likes368 downloads14d agoHugging Face21laion /terminal_bench_2_tasktrove_dq_taco_step15_30b_a3b_20260729_222705 Agent trace dataset OpenCode/Harbor rollout traces from the MarinSkyRL run rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260726-235656-574ba8, exported with make_and_upload_trace_dataset --episodes last (the last episode of each trial — the rollouts the policy was trained on). Coverage Built from the complete trial set on durable object storage, not from a local evidence bundle. quantity value trial directories on object storage 21711 trials with a… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_taco_step15_30b_a3b_20260729_222705.text10K<n<100K0 likes329 downloads2mo agoHugging Face22RandyHuynh5815 /TACO-Waste-Recognitionimage1K<n<10K1 likes314 downloads4y agoHugging Face23Myungkyu /RMBench-taco-luna RMBench-taco-luna RMBench training episodes (9 tasks, 449 episodes, 30 fps) with dense high-level labels produced by the TACOR offline annotator: GPT-5.6 Luna (gpt-5.6-luna) reads the frames of each episode sampled every 25 frames as labelled images and labels every sampled frame under the task-specific context (taco) induced for that task. Each tick carries the current subtask, the running textual memory and the visual-memory operations (keyframe store / retrieval) that the… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RMBench-taco-luna.videorobotics1K<n<10K1 likes306 downloads10d agoHugging Face24saaduddinM /OXE_taco_play_embeddingsLanguage Table (LeRobot) — Embedding-Only Release (DINOv3 + SigLIP2 image features; EmbeddingGemma task-text features) This repository packages a re-encoded variant of IPEC-COMMUNITY/taco_play_lerobot where raw videos are replaced by fixed-length image embeddings, and task strings are augmented with text embeddings. All indices, splits, and semantics remain consistent with the source dataset while storage and I/O are substantially lighter. To make the dataset practical to upload/download and… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/OXE_taco_play_embeddings.tabularrobotics100K<n<1M0 likes289 downloads1y agoHugging Face25Akanezora /TACO-Benchmark TACO-Benchmark TACO (Text-to-SQL with Ambiguous and Cross-database Open-domain queries) is a benchmark for real-world data-lake Text-to-SQL. 📢 News (2026): TACO has been accepted to VLDB 2026! 🎉📄 Paper: arXiv:2606.14201 GitHub (code & evaluation): Akanezora0/TACO-Benchmark Google Drive mirror: TACO-Benchmark.zip Overview Unlike Spider or BIRD — where the target database is known and schemas are clean — TACO evaluates systems on three challenges common in… See the full description on the dataset page: https://huggingface.co/datasets/Akanezora/TACO-Benchmark.text-generation10K<n<100K1 likes270 downloads3mo agoHugging Face26pictograph /taco View on Pictograph · Pictograph Research · Creative Commons Attribution 4.0 About TACO is a computer-vision dataset curated and annotated on Pictograph. The most common detected objects are grass, sidewalk, bush, bottle, frisbee, ruins. On Pictograph you can browse every annotated image, fork it into your own workspace in one click, export it in a dozen formats, or train a model on it directly. At a glance Metric Value Images 1,500 Annotations 4… See the full description on the dataset page: https://huggingface.co/datasets/pictograph/taco.imageobject-detection1K<n<10K1 likes259 downloads2mo agoHugging Face27gijs /tacos-captioningaudio10K<n<100K1 likes246 downloads1y agoHugging Face28rmcpantoja /taco2-checkpoints0 likes184 downloads2y agoHugging Face29aayushis196 /pi05-libero-taco-eval0 likes181 downloads15d agoHugging Face30Myungkyu /RoboDojo-taco-gemini RoboDojo-taco-gemini RoboDojo long-horizon episodes (8 tasks, 800 episodes, 25 fps) with dense high-level labels produced by the TACOR offline annotator: Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 25 frames) and labels every sampled frame under the task-specific context (taco) induced for that task, in its hybrid form: labels name the target object's image coordinates only where words cannot identify it (a random instance of a class that… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/RoboDojo-taco-gemini.videorobotics1K<n<10K0 likes173 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.