datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
humanego_serve_bread_lingbot_lerobot_with_latents
HumanEgo Serve Bread LingBot LeRobot With Latents
This dataset contains LeRobot-format robot demonstrations for the task:
pick up the bread and place it on the plate
The repository has two standalone LeRobot-style roots:
humanego_serve_bread_lingbot_eef_train: 55 episodes, 41,603 frames, 55 videos.
humanego_serve_bread_lingbot_eef_val: 6 episodes, 5,533 frames, 6 videos.
Each split includes:
data/: episode parquet files.
videos/: MP4 videos for observation.images.ego_rgb.… See the full description on the dataset page: https://huggingface.co/datasets/Coffeecoderss/humanego_serve_bread_lingbot_lerobot_with_latents.coffee-brew-cross-embodiment-rich-modality-sample
Cross-Embodiment Coffee Brew — Rich-Modality 10-Episode Inspection Sample
10 full-modality cross-embodiment coffee-brew episodes: 5 single-arm Franka Panda + 5 single-arm WidowXAI, 14,401 frames, 5 RGB views per robot, task-camera depth and segmentation, native robot state/action, end-effector trajectories, 6-DoF object poses, and QA annotations.
✅ Use it / ❌ Skip it
Use it for
Inspecting loaders, schemas, camera coverage, depth, segmentation, object poses… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/coffee-brew-cross-embodiment-rich-modality-sample.UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper
The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files.
Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeZongzi/UIIS.coffee-bags-tabular
Coffee Bags — Premium Pricing (Tabular)
34 retail coffee bags described by seven package attributes, with a binary target
for whether a bag is premium-priced per ounce. Built for 24-679.
Property
Value
Splits
train (345), validation (5), test (6)
Features
7
Targets
is_premium (binary), price_per_oz (continuous)
Balance (all originals)
17 premium / 17 standard
Purpose
Can package attributes alone predict whether a coffee is expensive per… See the full description on the dataset page: https://huggingface.co/datasets/ssg1/coffee-bags-tabular.difficulty-E2H-AMC-generations
Generations Dataset: E2H-AMC
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k
int
Number of samples generated
temperature
float
Sampling temperature
max_len
int
Maximum… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-E2H-AMC-generations.croppie_coffee_ugCroppie © 2024 by Producers Direct and Alliance Bioversity & CIAT is licensed under CC BY-SA 4.0
Funded by: Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ) Fair Forward Initiative - AI for All
Croppie training datasets
General information
Croppie dataset for machine-vision assisted coffee cherry detection. The dataset is made of a mix of Arabica and Robusta coffee tree parts (with and without a background isolation element) with individual bounding… See the full description on the dataset page: https://huggingface.co/datasets/rgautroncgiar/croppie_coffee_ug.robocasa_20260430T030150Z_full_run_sweeten_coffee ---
pretty_name: RoboCasa Trajectories Single
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
---
# RoboCasa Trajectories Single
This dataset contains one row per RoboCasa trajectory / episode.
## Structure
Each row is one trajectory / episode.
Episode-level JSON is stored inline:
adapted_trajectory
original_trajectory
execution_metadata
Step-level data is stored in aligned sequence columns:… See the full description on the dataset page: https://huggingface.co/datasets/DorianAtSchool/robocasa_20260430T030150Z_full_run_sweeten_coffee.COFFEE-Dataset-both-correct-and-incorrect-edit-feedback
Dataset Card for "COFFEE-Dataset-both-correct-and-incorrect-edit-feedback"
More Information needed
roasterdb-specialty-coffee-sample
☕ RoasterDB — Specialty Coffee Dataset (Free Sample)
Full dataset: roasterdb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A free sample of RoasterDB: a structured dataset of specialty-coffee products scraped from the direct storefronts of curated artisan roasters worldwide, with tasting notes normalized to the Specialty Coffee Association (SCA) Flavor Wheel and a source URL on every record so any fact can be re-verified.
This sample contains 100… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/roasterdb-specialty-coffee-sample.difficulty-gsm8k-generations
Generations Dataset: gsm8k
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k
int
Number of samples generated
temperature
float
Sampling temperature
max_len
int
Maximum… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-gsm8k-generations.coffee_rust_multispec_classification
Coffee Rust Multispec Classification
A dataset for image classification of Coffee Rust Multispec Classification. The dataset contains 1,120 images across 2 classes: NoRust, Rust.Images per class:
NoRust: 273
Rust: 847
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{arocatrujillo2025colombian,
title={Colombian coffee tree leaves multispectral images dataset},
author={Aroca-Trujillo, Jorge Luis… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/coffee_rust_multispec_classification.difficulty-aime_2025-generations
Generations Dataset: aime_2025
Paper: LLMs Encode Their Failures: Predicting Success from Pre-Generation ActivationsCode: GitHub
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-aime_2025-generations.difficulty-MATH-generations
Generations Dataset: MATH
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k
int
Number of samples generated
temperature
float
Sampling temperature
max_len
int
Maximum… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-MATH-generations.buzz_sources_090_coffeescriptpika-math-generations
PIKA MATH Generations Dataset
A comprehensive dataset of MATH problem solutions generated by different language models with various sampling parameters.
Paper (arXiv) | GitHub Repository
Dataset Description
This dataset contains code generation results from the MATH Dataset evaluated across multiple models. It was created to support the PIKA (Probe-Informed K-Aware Routing) project, which investigates how LLMs encode their own likelihood of success in their internal… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/pika-math-generations.COFFEE-Dataset
Dataset Card for "COFFEE-Dataset"
This is the official dataset for COFFEE: Boost Your Code LLMs by Fixing Bugs with Feedback
COFFEE dataset is built for training a critic that generates natural language feedback given an erroneous code.
Overall Filtered ratio: 12.65%
Short Feedback: 0.00% (0 samples)
stdin readline present: 1.37% (639 samples)
Low Diff Score: 7.79% (3622 samples)
Low Variable Overlap: 1.75% (813 samples)
Variable Name: 1.74% (807 samples)
The number of problem id in… See the full description on the dataset page: https://huggingface.co/datasets/LangAGI-Lab/COFFEE-Dataset.agungpambudi_trends-product-coffee-shop-sales-revenue-dataset
Maven Roasters: Coffee Shop Sales & Revenue Data
Unveiling Trends: Time Analysis, Transaction & Revenue in Coffee Shop Sales Data
Dataset Info
Source: Kaggle
Original Size: 2.54 MB
Kaggle Downloads: 4,364
Files: 2
Files
coffee-shop-sales-revenue.csv
coffee-shop-sales-revenue.parquet
Mirrored from Kaggle
coffeescript-code-suite
CoffeeScript Code Suite
CoffeeScript Code Suite is a public code dataset built from permissively licensed open-source CoffeeScript repositories. It is designed for three practical uses:
CoffeeScript domain adaptation and continued pretraining through raw_corpus examples.
CoffeeScript completion training through completion examples.
CoffeeScript and JavaScript translation training through coffee_to_js and js_to_coffee examples.
The dataset was assembled automatically from public… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/coffeescript-code-suite.mini-coffee-test-trajectories
Mini Coffee-Test Trajectories
200 episodes of "unfamiliar house → find kitchen → make coffee," solved by
a privileged BFS oracle from the
Mini Coffee-Test harness. I generated this mostly
as a byproduct of building the harness itself — once the oracle agent
existed, dumping its rollouts to a dataset was nearly free.
Each JSONL record is one full episode:
{
"seed": 0,
"house": {"entrance_id": ..., "kitchen_id": ..., "rooms": {...}},
"trajectory": [
{"obs": {...}… See the full description on the dataset page: https://huggingface.co/datasets/1rahu/mini-coffee-test-trajectories.CoffeeSales
Coffee Sales
This is a dataset copied from Kaggle. You can see the original dataset here: https://www.kaggle.com/datasets/ihelon/coffee-sales
The following is the original readme of this dataset:
About Dataset
Overview
This dataset contains detailed records of coffee sales from a vending machine.
The vending machine is the work of a dataset author who is committed to providing an open dataset to the community.
It is intended for analysis of purchasing patterns… See the full description on the dataset page: https://huggingface.co/datasets/tablegpt/CoffeeSales.coffeeMakingDemo
Coffee Making Demo (TsFile)
Apache TsFile version of allday-technology/coffeeMakingDemo.
Overview
A LeRobot (v2.1) teleoperation dataset of a coffee-making demonstration recorded on a Trossen AI stationary robot. Each episode is a single-arm (7-DoF) trajectory with synchronized proprioceptive state and action commands at 30 fps.
The dataset covers the following tasks:
pick up a coffee capsule
pick up an empty cup by the handle
Robot: trossen_ai_stationary… See the full description on the dataset page: https://huggingface.co/datasets/THULab/coffeeMakingDemo.known-actions-tracescoffee-order-zhtw
本資料集由OllaForge生成
Coffee Order Dataset (繁體中文/Traditional Chinese)
專為咖啡點餐場景設計的繁體中文多輪對話資料集,適用於訓練任務導向對話系統。
資料集描述
此資料集包含模擬咖啡店點餐場景的多輪對話,涵蓋各種點餐情境,包括:
基本點餐流程
訂單修改與取消
處理菜單外品項
處理超出限制的請求(如加兩份濃縮)
口語化表達理解
語言
繁體中文(台灣)
包含台灣口語表達(如「ㄋㄟㄋㄟ」)
資料規模
項目
數量
總對話數
2,939
平均輪數
2-6 輪
語言
繁體中文
資料格式
每筆資料為 JSONL 格式,包含 conversations 欄位:
{
"conversations": [
{
"role": "system",
"content":… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/coffee-order-zhtw.HeadlineHunter
HeadlineHunter
HeadlineHunter is a novel Document Layout Analysis Dataset centred on newspapers.
The newspapers currently in the dataset are from The Daily Monitor (Uganda), but we hope to add more as time progresses.
Class Labels
{0: 'Ad',
1: 'Table',
2: 'byline',
3: 'caption',
4: 'deck',
5: 'folio',
6: 'headline',
7: 'illustration',
8: 'jumpline',
9: 'masthead',
10: 'photograph',
11: 'story'}
Citation Information
@ONLINE{Headline Hunter… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/HeadlineHunter.coffee-pod-capsule-prices-raw-dataset-2026
Coffee Pod & Capsule Prices Raw Dataset (2026)
Analyze 151,252 product-level listed retail prices for single-serve coffee pods and capsules across 12 U.S. ZIP markets and every date from July 13 through August 10, 2026. The CSV preserves full product titles, listed package prices, resolved pod counts, and a comparable 24-pod package-equivalent price.
What “raw” means here: unaggregated, product-level retail price observations after quality filtering. The file also contains… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/coffee-pod-capsule-prices-raw-dataset-2026.coffee_detection
Coffee Detection
A dataset for object detection of coffee beans. The dataset contains 3,254 images with 126,840 bounding box annotations across 5 categories.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{sanya2024coffee,
title={Coffee and cashew nut dataset: A dataset for detection, classification, and yield estimation for machine learning applications},
author={Sanya, Rahman and Nabiryo, Ann… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/coffee_detection.instant-coffee-prices-raw-dataset-2026
119,960 raw U.S. instant coffee price observations across 12 ZIP markets and 29 days.
Instant Coffee Prices Raw Dataset (2026)
Analyze 119,960 unaggregated product-level listed retail prices for instant coffee powder and granules across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field.
What “raw” means here: unaggregated… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/instant-coffee-prices-raw-dataset-2026.mimicgen_sim_coffee_preparation_d0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 100,
"total_frames": 68161,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iantc104/mimicgen_sim_coffee_preparation_d0.ground-coffee-prices-raw-dataset-2026
134,333 raw U.S. ground coffee price observations across 12 ZIP markets and 29 days.
Ground Coffee Prices Raw Dataset (2026)
Analyze 134,333 unaggregated product-level listed retail prices for ground coffee across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field.
What “raw” means here: unaggregated product-level… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/ground-coffee-prices-raw-dataset-2026.COFFEE-Dataset
