datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coffee-bean-grading-dataset
An Image Dataset of Pre-Roast Robusta Coffee Beans with Polygon Annotations for Automated Grading
Abstract
Automated quality assessment of raw agricultural products is critical for ensuring fair trade and supply chain efficiency. This dataset presents 3,877 high-resolution images of pre-roast Arabica coffee beans, collected from farms in **Coorg, Karnataka (India)**—a major coffee-producing region. Each bean is categorized into one of four quality grades:
Grade… See the full description on the dataset page: https://huggingface.co/datasets/SamruddhK/coffee-bean-grading-dataset.Coffee_leaves_sections_FO
Dataset Card for coffee_leaves_anomalib_2
This is a FiftyOne dataset with 35962 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("pjramg/Coffee_leaves_sections_FO")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/pjramg/Coffee_leaves_sections_FO.mimicgen_coffee_preparation_d1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1000,
"total_frames": 224403,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mishmish66/mimicgen_coffee_preparation_d1.robocasa-heldout-coffeeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1023,
"total_frames": 295793,
"total_tasks": 236,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1023"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Fascetta/robocasa-heldout-coffee.coffee-bean-grading-dataset
An Image Dataset of Pre-Roast Robusta Coffee Beans with Polygon Annotations for Automated Grading
Abstract
Automated quality assessment of raw agricultural products is critical for ensuring fair trade and supply chain efficiency. This dataset presents 3,877 high-resolution images of pre-roast Arabica coffee beans, collected from farms in **Coorg, Karnataka (India)**—a major coffee-producing region. Each bean is categorized into one of four quality grades:… See the full description on the dataset page: https://huggingface.co/datasets/kushi-ai-2027/coffee-bean-grading-dataset.Coffee-Machine-Parts-Detection
Coffee Machine Parts Detection
Object detection dataset for coffee machines and their parts.
Contents
train
validation
test
Labels
group_head
steam_knob
coffee_spouts
milk_reservoir
portafilter
bean_hopper
steam_wand
filter_basket
drip_tray
control_panel
carafe
coffee_machine
power_switch
water_reservoir
Format
COCO source data
Hugging Face imagefolder package with metadata.jsonl
Bounding boxes for object detection
license: mit… See the full description on the dataset page: https://huggingface.co/datasets/start2fix/Coffee-Machine-Parts-Detection.robocasa-kitchen-coffeeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 162,
"total_frames": 39008,
"total_tasks": 3,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:162"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Fascetta/robocasa-kitchen-coffee.robocasa_kitchen_coffeeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 162,
"total_frames": 39008,
"total_tasks": 3,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:162"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iAyoD/robocasa_kitchen_coffee.coffee-lamp
Synthetic Image-Classification Dataset
Synthetic image-classification dataset generated with stable diffusion
(zerogpu_sdxl_turbo) using text-to-image from class names + short descriptions.
Classes
Label
Images
background
20
coffee-mug
20
lamp
20
Layout
train/<label>/<label>.<id>.jpg
test/<label>/<label>.<id>.jpg
metadata.csv
Loading
from datasets import load_dataset
ds = load_dataset("imagefolder"… See the full description on the dataset page: https://huggingface.co/datasets/edgeimpulse/coffee-lamp.UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper
The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files.
Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeZongzi/UIIS.arabica_coffee_leaf_disease_classification
Arabica Coffee Leaf Disease Classification
A dataset for disease classification of Arabica Coffee Leaf. The dataset contains 58,549 images across 5 classes: Cerscospora, Healthy, Leaf_rust, Miner, Phoma.Images per class:
Cerscospora: 7,681
Healthy: 18,983
Leaf_rust: 8,336
Miner: 16,978
Phoma: 6,571
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{jepkoech2021arabica,
title={Arabica coffee leaf… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/arabica_coffee_leaf_disease_classification.croppie_coffee_ugCroppie © 2024 by Producers Direct and Alliance Bioversity & CIAT is licensed under CC BY-SA 4.0
Funded by: Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ) Fair Forward Initiative - AI for All
Croppie training datasets
General information
Croppie dataset for machine-vision assisted coffee cherry detection. The dataset is made of a mix of Arabica and Robusta coffee tree parts (with and without a background isolation element) with individual bounding… See the full description on the dataset page: https://huggingface.co/datasets/rgautroncgiar/croppie_coffee_ug.robocasa_20260430T030150Z_full_run_sweeten_coffee ---
pretty_name: RoboCasa Trajectories Single
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
---
# RoboCasa Trajectories Single
This dataset contains one row per RoboCasa trajectory / episode.
## Structure
Each row is one trajectory / episode.
Episode-level JSON is stored inline:
adapted_trajectory
original_trajectory
execution_metadata
Step-level data is stored in aligned sequence columns:… See the full description on the dataset page: https://huggingface.co/datasets/DorianAtSchool/robocasa_20260430T030150Z_full_run_sweeten_coffee.mimicgen_coffee_d0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1000,
"total_frames": 223130,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mishmish66/mimicgen_coffee_d0.coffee_rust_multispec_classification
Coffee Rust Multispec Classification
A dataset for image classification of Coffee Rust Multispec Classification. The dataset contains 1,120 images across 2 classes: NoRust, Rust.Images per class:
NoRust: 273
Rust: 847
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{arocatrujillo2025colombian,
title={Colombian coffee tree leaves multispectral images dataset},
author={Aroca-Trujillo, Jorge Luis… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/coffee_rust_multispec_classification.Coffee-Bean-Package-Images-OCR
Dataset Description: This dataset contains comprehensive coffee bean package images specifically designed for Optical Character Recognition (OCR) tasks.
The collection includes three distinct datasets: LRCB (Low-Resolution Coffee Bean), HRCB (High-Resolution Coffee Bean), and POIE (Product OCR Image Evaluation), providing researchers and practitioners with diverse image qualities and real-world scenarios for developing and evaluating OCR systems.
Also, we provide antonation datasets… See the full description on the dataset page: https://huggingface.co/datasets/Thi-Thu-Huong/Coffee-Bean-Package-Images-OCR.RoCoLe-Coffeehttps://www.sciencedirect.com/science/article/pii/S2352340919307693
license: mit
HeadlineHunter
HeadlineHunter
HeadlineHunter is a novel Document Layout Analysis Dataset centred on newspapers.
The newspapers currently in the dataset are from The Daily Monitor (Uganda), but we hope to add more as time progresses.
Class Labels
{0: 'Ad',
1: 'Table',
2: 'byline',
3: 'caption',
4: 'deck',
5: 'folio',
6: 'headline',
7: 'illustration',
8: 'jumpline',
9: 'masthead',
10: 'photograph',
11: 'story'}
Citation Information
@ONLINE{Headline Hunter… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/HeadlineHunter.Brazilian_Coffee_Scenes
Dataset Card for "Brazilian_Coffee_Scenes"
Licensing Information
[CC BY-NC]
Citation Information
Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?
@inproceedings{penatti2015deep,
title = {Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?},
author = {Penatti, Ot{\'a}vio AB and Nogueira, Keiller and Dos Santos, Jefersson A},
year = 2015… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/Brazilian_Coffee_Scenes.coffee_detection
Coffee Detection
A dataset for object detection of coffee beans. The dataset contains 3,254 images with 126,840 bounding box annotations across 5 categories.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{sanya2024coffee,
title={Coffee and cashew nut dataset: A dataset for detection, classification, and yield estimation for machine learning applications},
author={Sanya, Rahman and Nabiryo, Ann… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/coffee_detection.mimicgen_coffee_d1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1000,
"total_frames": 224403,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mishmish66/mimicgen_coffee_d1.Coffee_leaves_rocole_FO
Dataset Card for coffee_rocole_patches
This is a FiftyOne dataset with 2329 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("pjramg/Coffee_leaves_rocole_FO")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/pjramg/Coffee_leaves_rocole_FO.v122_coffee_pod_sade_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 23,
"total_frames": 6537,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:23"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/dageorge1111/v122_coffee_pod_sade_test.v122_coffee_pod_sade_trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 90,
"total_frames": 23744,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/dageorge1111/v122_coffee_pod_sade_train.v200_coffee_pod_traincoffee-beans
Dataset Card for Beans
Dataset Summary
Coffee Beans Grading
Supported Tasks and Leaderboards
image-classification: Based on a coffee bean grading, the goal of this task is to grade single beans for clusterization.
Languages
Indonesia
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/rasyidf/coffee-beans.coffee-beanscolombian_coffeecoffee_bean_quality_classification
Coffee Bean Quality Classification
A dataset for quality classification of Coffee Beans. The dataset contains 464 images across 9 classes: A, AA, AAA, AB, Bits, Bulk, C, PB-I, PB-II.
Images per class:
A: 50
AA: 61
AAA: 50
AB: 50
Bits: 51
Bulk: 50
C: 51
PB-I: 51
PB-II: 50
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{bj2025cbd,
title={CBD: Coffee Beans Dataset},
author={BJ, Bipin Nair and KM… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/coffee_bean_quality_classification.coffee_rocole
Dataset Card for coffee_rocole
This is a FiftyOne dataset with 1560 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("pjramg/coffee_rocole")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/pjramg/coffee_rocole.
