datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.BEANS-Zero
BEANS-Zero
Version: 0.1.0
Created on: 2025-04-12
Creators:
Earth Species Project (https://www.earthspecies.org)
Overview
BEANS-Zero is a bioacoustics benchmark designed to evaluate multimodal audio-language models in zero-shot settings. Introduced in the paper NatureLM-audio paper (Robinson et al., 2025), it brings together tasks from both existing datasets and newly curated resources.
The benchmark focuses on models that take a bioacoustic audio input (e.g., bird or… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/BEANS-Zero.AI2_Alphabot_2_pour_mung_beans
AI2_Alphabot_2_pour_mung_beans
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 213
Total Frames: 181968
FPS: 30
Dataset Size: 7.47 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_pour_mung_beans.Agilex_Cobot_Magic_sweep_coffee_beans
Agilex_Cobot_Magic_sweep_coffee_beans
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 359
Total Frames: 331666
FPS: 30
Dataset Size: 13.24 GB
Robot Name: Agilex_Cobot_Magic
End-Effector Type: two_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_sweep_coffee_beans.RMC-AIDA-L_pour_coffee_beans
RMC-AIDA-L_pour_coffee_beans
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
kitchen
restaurant
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
takeout
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_pour_coffee_beans.Agilex_Cobot_Magic_scoop_coffee_beans
Agilex_Cobot_Magic_scoop_coffee_beans
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Agilex_Cobot_Magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
scoop
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_scoop_coffee_beans.jev-stage2-image-beans-pilot
Beans: one natural question per image
Open the corrected preview.
natural_v4 is the recommended and default preview: 100 original images, 100 rows, one three-way condition-class Choice question per image. All targets come directly from the source labels column (34 angular leaf spot, 33 bean rust, 33 healthy). Original image bytes and source annotations are unchanged.
Example question: “Which source-defined condition class describes the bean leaf?” Options: angular_leaf_spot… See the full description on the dataset page: https://huggingface.co/datasets/FaroukMoc2/jev-stage2-image-beans-pilot.beans_watkins
Dataset Card for "beans_watkins"
Dataset Description
Paper:
https://doi.org/10.1121/2.0000358
Dataset Summary
This dataset contains annotated recordings of marine mammal sounds with splits and preprocessing like described in BEANS. It is used for classification tasks.
Data Splits
train
train_low
valid
test
1017
203
339
339
autotrain-data-coffee-beans
AutoTrain Dataset for project: coffee-beans
Dataset Description
This dataset has been automatically processed by AutoTrain for project coffee-beans.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<224x224 RGB PIL image>",
"feat_width": 224,
"feat_height": 224,
"target": 1,
"feat_xmin": 22,
"feat_ymin": 61… See the full description on the dataset page: https://huggingface.co/datasets/everycoffee/autotrain-data-coffee-beans.BEANS-Next
BEANS-Next
Audio files live under audio/. Metadata is in metadata.parquet with a
file_name column (repo-relative paths) so the Hugging Face Dataset Viewer can
play clips—only file_name uses that convention; tier 4 rows instead use
context_audio_paths (list) and query_audio_path so the viewer is not confused
by several *_file_name-like columns. Column task is the benchmark task id
(same strings as the old subset column in other layouts). Column tier is an
integer in {1,2,3,4}… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/BEANS-Next.beans_rfcx
Dataset Card for "beans_rfcx"
Dataset Description
Paper: https://doi.org/10.1016/j.ecoinf.2020.101113
Dataset Summary
This dataset contains continuous soundscape
recordings of 24 species of frogs and birds collected by Rain-
forest Connection (RFCx) with splits and preprocessing like described in BEANS. It is used for detection tasks.
Data Splits
train
train_low
valid
test
2836
964
945
946
beans_bats
Dataset Card for "beans_bats"
Dataset Description
Paper: https://doi.org/10.1038/sdata.2017.143
Dataset Summary
This dataset contains annotated recordings of Egyptian fruit bats with splits and preprocessing like described in BEANS. It is used for classification tasks.
Data Splits
train
train_low
valid
test
6000
1200
2000
2000
Regression_cocoa_beansbeansBeans is a dataset of images of beans taken in the field using smartphone
cameras. It consists of 3 classes: 2 disease classes and the healthy class.
Diseases depicted include Angular Leaf Spot and Bean Rust. Data was annotated
by experts from the National Crops Resources Research Institute (NaCRRI) in
Uganda and collected by the Makerere AI research lab.spotlight-beans-enrichment
Dataset Card for "spotlight-beans-enrichment"
More Information needed
beans_dcase
Dataset Card for "beans_dcase"
Dataset Description
Paper: https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Morfi_52.pdf
Dataset Summary
This is the dataset used for the DCASE 2021 Task and contains annotated mammal and bird multi-species recordings with splits and preprocessing like described in BEANS. It is used for detection tasks.
Data Splits
train
train_low
valid
test
702
151
234
232
beans_dogs
Dataset Card for "beans_dogs"
Dataset Description
Paper: https://doi.org/10.1016/j.anbehav.2003.07.016
Dataset Summary
This dataset contains annotated recordings of domestic dog barks with splits and preprocessing like described in BEANS. It is used for classification tasks.
Data Splits
train
train_low
valid
test
415
83
139
139
beans_enabirds
Dataset Card for "beans_enabirds"
Dataset Description
Paper: https://doi.org/10.1002/ecy.3329
Dataset Summary
This dataset contains annotated recordings of Eastern North American birds from the dawn chorus with splits and preprocessing like described in BEANS. It is used for detection tasks.
Data Splits
train
train_low
valid
test
231
77
77
77
beans_sc
Dataset Card for "beans_sc"
Dataset Description
Paper: https://doi.org/10.48550/arXiv.1804.03209
Dataset Summary
This dataset contains single-word utterances covering 35 English words including digits, commands, and others, spoken by multiple speakers with splits and preprocessing like described in BEANS. It is used for classification tasks and is an auxiliary dataset in BEANS.
Data Splits
train
train_low
valid
test
84843
849
9981
11005… See the full description on the dataset page: https://huggingface.co/datasets/DBD-research-group/beans_sc.beans_cbi
Dataset Card for "beans_cbi"
Dataset Description
Paper: https://www.kaggle.com/competitions/birdsong-recognition/overview
Dataset Summary
The original recordings from this dataset are from the Cornell Bird Identification competition hosted on Kaggle. The training set consists of bird recordings uploaded to xeno-canto3 by volunteer
users with splits and preprocessing like described in BEANS. It is used for classification tasks.
Data Splits
train… See the full description on the dataset page: https://huggingface.co/datasets/DBD-research-group/beans_cbi.camel-beanspiper_umi_grind_beansThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "piperx_bimanual_eef_6d",
"total_episodes": 59,
"total_frames": 51459,
"total_tasks": 1,
"total_videos": 118,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:59"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/axiboai/piper_umi_grind_beans.piper_pour_beans_jointThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "piperx_bimanual_joint",
"total_episodes": 45,
"total_frames": 41191,
"total_tasks": 1,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:45"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/axiboai/piper_pour_beans_joint.beans-mini
AnnotateIt · Open the app · Models & datasets · Documentation
AnnotateIt Beans Mini
Small, deterministic, AnnotateIt-compatible samples derived from Beans.
Upstream revision: 27aa014ce09b193e1a6f58112d4a66e0eddb69c5Upstream license: MIT
File
Sample
Tasks
Format
Images
Size
SHA-256
Derived task
beans-classification-mini.zip
Beans Classification Mini
Classification
Datumaro
45
5.84 MiB
5dc1b75552f96b67d6843ca7cb24cc7b431e3789b63158df4786b4ab8bcc0637
no… See the full description on the dataset page: https://huggingface.co/datasets/AnnotateIt/beans-mini.pour_coffee_beansThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 21,
"total_frames": 25239,
"total_tasks": 1,
"total_videos": 21,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jcoleharrison/pour_coffee_beans.beans-outlier
Dataset Card for "beans-outlier"
📚 This dataset is an enhancved version of the ibean project of the AIR lab.
The workflow is described in the medium article: Changes of Embeddings during Fine-Tuning of Transformers.
Explore the Dataset
The open source data curation tool Renumics Spotlight allows you to explorer this dataset. You can find a Hugging Face Space running Spotlight with this dataset here: https://huggingface.co/spaces/renumics/beans-outlier
Or you can… See the full description on the dataset page: https://huggingface.co/datasets/renumics/beans-outlier.beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/javieratorres/beans.beans_humbugdb
Dataset Card for "beans_humbugdb"
Dataset Description
Paper: https://doi.org/10.48550/arXiv.2110.07607
Dataset Summary
This dataset contains annotated recordings of wild and cultured mosquito wingbeat sounds with splits and preprocessing like described in BEANS. It is used for classification tasks.
Data Splits
train
train_low
valid
test
5577
1115
1859
1859
beans_hiceas
Dataset Card for "beans_hiceas"
Dataset Description
Paper: https://doi.org/10.25921/e12p-gj65
Dataset Summary
This dataset contains a subset of pas-
sive acoustic data collected using a multi-channel towed hy-
drophone during the Hawaiian Islands Cetacean and Ecosys-
tem Assessment Survey (HICEAS) in 2017 with splits and preprocessing like described in BEANS. It is used for detection tasks.
Data Splits
train
train_low
valid
test
407
No… See the full description on the dataset page: https://huggingface.co/datasets/DBD-research-group/beans_hiceas.coffee-beans
