datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
COCOANet
COCOANet
A high-fidelity CFD dataset of parametric (CCA) Aircraft geometries and flight condition for aerodynamic surrogate modelling and Aerodynamic Shape Optimization, created as part of ShapeBench.
Dataset Summary
Geometries
401
CFD runs
3,570
Design parameters
16 geometric
Flight parameters
3 (angle of attack, velocity, altitude)
CAD Kernal
nTop
Solver
Flow360
Sampling
Latin Hypercube Sampling (LHS), seed 42
Files… See the full description on the dataset page: https://huggingface.co/datasets/ShapeBench/COCOANet.CoCoaSpec
CoCoaSpec: A Multimodal hyperspectral dataset of cocoa beans with physicochemical annotation
Overview
The CoCoaSpec dataset is a multimodal hyperspectral imaging dataset of Colombian cocoa beans with detailed physicochemical annotations.It was created to support research on non-destructive cocoa quality assessment, spectral data analysis, and multimodal data fusion.
The dataset includes hyperspectral images acquired with four different devices, along with reference… See the full description on the dataset page: https://huggingface.co/datasets/ecos-nord-ginp-uis/CoCoaSpec.CoCoaSpec2amini_cocoa_coco_datasetCOCO_AI
VISUAL COUNTER TURING TEST - COCO DATASET
The Visual Counter Turing Test (VCT²) dataset is introduced in the paper“Visual Counter Turing Test (VCT²): Discovering the Challenges for AI-Generated Image Detection and Introducing Visual AI Index (V_AI)”,accepted at IJCNLP–AACL 2025 and available on arXiv:2411.16754.
This dataset aims to benchmark and analyze the challenges of AI-generated image detection (AGID) in the era of advanced text-to-image models.It provides a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/NasrinImp/COCO_AI.cocoa_agroforestry_multispectral
Cocoa Agroforestry Multispectral
An unlabeled image dataset of Cocoa Agroforestry Multispectral. The dataset contains 1,272 images with no classification, segmentation, or bounding-box annotations.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{lammoglia2024high,
title={High-resolution multispectral and RGB dataset from UAV surveys of ten cocoa agroforestry typologies in Côte d'Ivoire}… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/cocoa_agroforestry_multispectral.Mocha-trajectories
Mocha Trajectories Mini/Miniswe Subset
This dataset mirrors the mini/miniswe trajectory splits from ZeonLap/Mocha-trajectories, excluding the Kwai-Klear mini SWE-agent split.
Splits
Split
Trajectories
swe_rebench_dpskv32_miniswe_4k
4,389
swe_rebench_qwen3coder480b_mini_11k
12,789
swe_rebench_qwen3codernext_mini_17k
32,338
swe_smith_py_qwen3codernext_mini_7k
9,117
miniswe_kimik25_smith_pass16
115,965
Total
174,598
Citation
If you… See the full description on the dataset page: https://huggingface.co/datasets/cocoa-org/Mocha-trajectories.COCO_AIICOCO_ARC
Vision-Language Instruction Tuning: A Review and Analysis
Chen Li1, Yixiao Ge1, Dian Li2, and Ying Shan1.
1ARC Lab, Tencent PCG
2Foundation Technology Center, Tencent PCG
This paper is a review of all the works related to vision-language instruction tuning (VLIT). We will periodically update the recent public VLIT dataset and the VLIT data constructed by the pipeline in this paper.
📆 Schedule
Release New Vision-Language Instruction Data (periodically)… See the full description on the dataset page: https://huggingface.co/datasets/lllchenlll/COCO_ARC.CoralSpec-30MThis is a mirror copy of the original CoralSpec-30M dataset at KAUST library.
Please navigate to the "Files and versions" tab to download data.
The data is also available at Zenodo.
Note: the "coral_healthy_tip" label is currently unreliable. Please avoid using it.
Citation
If you find this dataset useful in your research, please cite:
@article{https://doi.org/10.1002/lol2.70156,
author = {Kang, Kaizhang and Heidrich, Wolfgang},
title = {CoralSpec-30M: A large-scale coral… See the full description on the dataset page: https://huggingface.co/datasets/cocoakang/CoralSpec-30M.CocoaMFDB_detection
CocoaMFDB Detection Object Detection
A dataset for detection of cocoa pods. The dataset contains 1,254 images with 1,482 bounding box annotations across 1 category.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{ayikpa2023cocoamfdb,
title={CocoaMFDB: A dataset of cocoa pod maturity and families in an uncontrolled environment in C{\^o}te d'Ivoire},
author={Ayikpa, Kacoutchy Jean and Mamadou… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/CocoaMFDB_detection.Regression_cocoa_beansCocoaMiningDS
Introduction
Agriculture remains a cornerstone of Ghana's economy. More than half of the land area is dedicated to agriculture, and its expansion continues to accelerate. Despite this growth, the sector faces compounding challenges, including climate-induced risks such as prolonged droughts, which aggravate low yields and threaten long-term productivity. A dimension of Ghana's agricultural landscape is agro-ecological regions, which not only sustain primary agricultural activities… See the full description on the dataset page: https://huggingface.co/datasets/ellaampy/CocoaMiningDS.coco-arvqa
COCO-ARVQA: Arabic Visual Question Answering over COCO 2017
Dataset Summary
COCO-ARVQA is an Arabic Visual Question Answering dataset built over images from MS COCO 2017 train2017.It provides Arabic questions, Arabic answers, answer lists, question identifiers, image identifiers, and COCO image file names.
This repository does not redistribute COCO images. Both the training and validation splits reference images from the official COCO 2017 train2017.zip archive.
Official… See the full description on the dataset page: https://huggingface.co/datasets/MouaffakAyoub/coco-arvqa.COCOACOCOA dataset targets amodal segmentation, which aims to recognize and segment objects beyond their visible parts. This dataset includes labels not only for the visible parts of objects, but also for their occluded parts hidden by other objects. This enables learning to understand the full shape and position of objects.common_voice_13_0_hi_pseudo_labelledcocoa_monilia_detection
Cocoa Monilia Detection
This dataset contains real RGB images of cocoa plants exhibiting monilia disease symptoms, collected in field environments across two Colombian locations. Images were captured using a variety of handheld smartphones during the September 2024 to February 2025 collection period, providing diverse natural lighting and environmental conditions for agricultural disease detection research. The dataset contains 1,950 images with 2,159 bounding box annotations… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/cocoa_monilia_detection.COCO-AB
General Information
Title: COCO-AB
Description:
The COCO-AB dataset is an extension of the COCO 2014 training set, enriched with additional annotation byproducts (AB).
The data includes 82,765 reannotated images from the original COCO 2014 training set.
It has relevance in computer vision, specifically in object detection and location.
The aim of the dataset is to provide a richer understanding of the images (without extra costs) by recording additional actions and interactions… See the full description on the dataset page: https://huggingface.co/datasets/coallaoh/COCO-AB.common_voice_13_0_zh_pseudo_labelledcocoa2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 10600,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yoshikokulala/cocoa2.asia-owid-cocoa-beans-production-by-region
Cocoa Beans Production By Region | Asia (Our World in Data)
🌏 435 observations · 8 Asia countries · 1961–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 435 observations of Cocoa Beans Production By Region data across 8 Asia countries, spanning 1961–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Cocoa Beans Production By Region
Geographic coverage
8… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-cocoa-beans-production-by-region.asia-owid-cocoa-bean-production
Cocoa Bean Production | Asia (Our World in Data)
🌏 435 observations · 8 Asia countries · 1961–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 435 observations of Cocoa Bean Production data across 8 Asia countries, spanning 1961–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Cocoa Bean Production
Geographic coverage
8 Asia countries · top rows shown… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-cocoa-bean-production.ProSR-Datacocoa_move3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so100_follower",
"total_episodes": 10,
"total_frames": 3457,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yoshikokulala/cocoa_move3.cocoa_move2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so100_follower",
"total_episodes": 10,
"total_frames": 3458,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yoshikokulala/cocoa_move2.cocoa_moveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so100_follower",
"total_episodes": 10,
"total_frames": 3454,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yoshikokulala/cocoa_move.Coffee-Cocoa-Diseasesvlm-dataset-cocoaasia-owid-cocoa-bean-yields
Cocoa Bean Yields | Asia (Our World in Data)
🌏 419 observations · 8 Asia countries · 1961–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 419 observations of Cocoa Bean Yields data across 8 Asia countries, spanning 1961–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Cocoa Bean Yields
Geographic coverage
8 Asia countries · top rows shown below, sorted… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-cocoa-bean-yields.cocoa_nikke
Dataset of cocoa/ココア/可可/코코아 (Nikke: Goddess of Victory)
This is the dataset of cocoa/ココア/可可/코코아 (Nikke: Goddess of Victory), containing 18 images and their tags.
The core tags of this character are bow, maid_headdress, twintails, long_hair, black_bow, hair_bow, pink_hair, bangs, brown_eyes, breasts, ahoge, hair_between_eyes, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/cocoa_nikke.
