datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brand-assetsartem-fold-towelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel.artem-pour-waterThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-pour-water.GLM-5.3-Flash-BF16-Teacher-Logits
GLM-5.3-Flash BF16 teacher logits
This dataset contains full-vocabulary float32 teacher logits from the immutable
zai-org/GLM-5.3-Flash-BF16 revision a6c167b62691b2bac901344b65cb651a70f53e43.
It keeps the sealed final KLD panel qualification-only and publishes the
separate non-final calibration panel under role-specific paths.
Qualification-only final windows: 25
Qualification-only final prediction positions: 51175
Vocabulary size: 154880
Teacher receipt:… See the full description on the dataset page: https://huggingface.co/datasets/brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits.brand-logosAn Ongoing SVG Collection of Many Multiples of Brand Logos
Object looks like this:
Shape:
├── brandName ""
├── brandWebsite ""
├── brandPresence[{
│ └── platform
│ └── url
│ └── username}]
├── brandLogo[{
│ └── fileName
│ └── svgPath
│ └── svgData
│ ├── meta
│ ├── width
│ ├── height
│ ├── viewbox
│ └── fill
│ └── svgRaw}]
└── brandColors[{
└── meta
├── primary
├── secondary
├── tertiary
├── quaternary
├──… See the full description on the dataset page: https://huggingface.co/datasets/mattrichmo/brand-logos.brand-assetsstock-market-data-warehouseVerdictAIfold-towel-3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/fold-towel-3.fold-towel-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/fold-towel-2.diffusion_policy_robocasa_activations_latest_chkpt
Diffusion Policy — RoboCasa Activations (latest checkpoint)
Per-step, per-episode activation traces collected from a DiffusionTransformerHybridImagePolicy (the diffusion_policy library) rolled out on RoboCasa benchmark tasks. Captured with collect_activations_robocasa.py at the latest training checkpoint.
These traces are the input expected by the conceptor / SAE steering pipelines under diffusion_policy/experiments/robocasa_steering/ and diffusion_policy/experiments/sae/ — see… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/diffusion_policy_robocasa_activations_latest_chkpt.GLM-5.3-BF16-full-logitsdual-lidar-umi-relativeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
12
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi2_x"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi-relative.metaworld_ml45-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "metaworld",
"total_episodes": 4391,
"total_frames": 358997,
"total_tasks": 44,
"total_videos": 0,
"total_chunks": 5,
"chunks_size": 1000,
"fps": 80,
"splits": {
"train": "0:4391"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/metaworld_ml45-v2.fold-towelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/fold-towel.P-KNN
The P-KNN framework source code is licensed under the MIT License. However, these precomputed scores are derived from dbNSFP.
Users are strictly bound by the dbNSFP licensing terms. For commercial use, you must obtain a commercial license directly from dbNSFP.
P-KNN Precomputed Scores Dataset
This dataset provides precomputed pathogenicity prediction scores generated by the P-KNN method using dbNSFP v5.2 (academic or commercial version) with joint… See the full description on the dataset page: https://huggingface.co/datasets/brandeslab/P-KNN.shapleymcg-qwen3-30b-a3b-reproducibility
ShapleyMCG Qwen3-30B-A3B reproducibility artifacts
This dataset preserves the calibration statistics, exact corrected-R10
EXL3/MCG candidates, source and corpus identities, BF16 teacher/student logits,
tokenwise KLD, allocations, attribution ledgers, hashes, and publication
receipts for the Qwen3-30B-A3B experiments in
brandonmmusic-max/shapleymcg.
The complete cross-checkpoint
results ledger
and
method specification
distinguish the predecessor routed-p2 allocator from the full… See the full description on the dataset page: https://huggingface.co/datasets/brandonmusic/shapleymcg-qwen3-30b-a3b-reproducibility.solar-pv-detection-brandenburg-dataset
Solar PV Ground Truth Dataset — Brandenburg Orthophotos
Hand-corrected ground-truth masks for photovoltaic detection on 20 cm GSD
aerial orthophotos (1 km × 1 km tiles, 4-band RGBI) of Brandenburg, Germany.
Companion to the model repository
solar-pv-segmentation-brandenburg.
Contents
Folder
Contents
gt_masks_selected/
1432 hand-corrected patch masks (256×256 px, uint8, 0=background / 1=PV), split into training/validation/testing
gt_masks_full_tiles/… See the full description on the dataset page: https://huggingface.co/datasets/Muemmel/solar-pv-detection-brandenburg-dataset.umi-fold-shirtThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/umi-fold-shirt.artem-fold-towel-filtered
Artem fold-towel filtered trajectories
Observation-only LeRobot v3 derivative of brandonyang/artem-fold-towel. It contains 781 demonstrations (1048134 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 14-D observation.state contains the smoothed, trajectory-optimized YAM-achievable UMI1 pose, normalized UMI1 gripper, UMI2 pose, and normalized UMI2 gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved; action is… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel-filtered.dual-lidar-umiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
12
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi2_x"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi.360-1M-Video-Part1droid_overlay_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 5289,
"total_frames": 862822,
"total_tasks": 3076,
"total_videos": 0,
"total_chunks": 6,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:5289"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/droid_overlay_v2.Plant-Diseases-PlantVillage-Dataset
Dataset Card for Dataset Name
Train and Test (20%) splits for the PlantVillage-Dataset on the subject of plant disease
@article{Mohanty_Hughes_Salathé_2016,
title={Using deep learning for image-based plant disease detection},
volume={7},
DOI={10.3389/fpls.2016.01419},
journal={Frontiers in Plant Science},
author={Mohanty, Sharada P. and Hughes, David P. and Salathé, Marcel},
year={2016},
month={Sep}}
Dataset Details
Dataset Description
Curated by: [More… See the full description on the dataset page: https://huggingface.co/datasets/BrandonFors/Plant-Diseases-PlantVillage-Dataset.brand-guidelines-pdfs
BrandGuide / Brand Genome Dataset
This dataset contains extracted brand guidelines for brand-level feature engineering and multimodal brand analysis.
Overview
BrandGuide is a large-scale collection of brand guidelines paired with corresponding brand assets. According to the paper, it covers 2,683 brands, spans 80 sectors, 103 regions, and 28 languages, with roughly 1M images and text assets collected over 2014–2025. The dataset is designed for interpretable ML and… See the full description on the dataset page: https://huggingface.co/datasets/brand-genome/brand-guidelines-pdfs.droid_overlay_fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 3776,
"total_frames": 995769,
"total_tasks": 3010,
"total_videos": 0,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:3776"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/droid_overlay_full.dual-lidar-combined-filtered-long-gripper
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 182 demonstrations (179951 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 12-D observation.state contains the smoothed, trajectory-optimized YAM-achievable path in the zero-origin UMI Cartesian convention. Raw UMI gripper widths remain as separate observations. The two original UMI… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered-long-gripper.droid_2000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1,
"total_frames": 167,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/droid_2000.360-1M-Video-Part2robot-umi-lidar3d-e2eThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
1200,
1920,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/robot-umi-lidar3d-e2e.
