datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kitti-yolo11n-robustness-benchmark
KITTI YOLO11n Robustness & Adversarial Benchmark Suite
This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels.
?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n)
Clean Baseline AP50: 0.3555
Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.lighting-invariant-bedroom-perception-robustness-benchmark
Lighting-Invariant Bedroom Perception & Robustness Benchmark
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.tokamark-robustness-data
TokaMark Sensor Robustness Benchmark Data
Associated paper: Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMarkAuthor: Neerav GuptaCode: github.com/Neerav-Gupta/tokamark-robustness
Dataset Description
This dataset contains pre-processed numpy arrays, trained model checkpoints, and experiment results from the first systematic robustness benchmark of plasma diagnostic ML models under realistic sensor failure, using the… See the full description on the dataset page: https://huggingface.co/datasets/Neerav-Gupta/tokamark-robustness-data.bangla-noise-robustness-dataeval_robustness_e9_3_full_dp_local_180k_fThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 175,
"total_frames": 73950,
"total_tasks": 1,
"total_videos": 350,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:175"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e9_3_full_dp_local_180k_f.eval_robustness_e6_3_100kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 24627,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e6_3_100k.eval_robustness_e6_3_260kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 22666,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e6_3_260k.eval_robustness_e6_3_200kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 23346,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e6_3_200k.Multimodal-Robustness-Benchmarkdecoding-robustness-results
Decoding Robustness Results
Mechanistic robustness evaluation results for language models under six input
perturbations: character replacement, BPE-token replacement, word replacement,
local token shuffle, typographical corruption, and synonym replacement.
The repository is organized by model and perturbation:
models/<model>/<perturbation>/<percentage>/evals.csv
The qwen2.5_1.5b/adversarial directory contains the separate adversarial
evaluation outputs and manifest. Failed or… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/decoding-robustness-results.code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.eval_robustness_e9_3_full_fk_local_240kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 177,
"total_frames": 74419,
"total_tasks": 1,
"total_videos": 354,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:177"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e9_3_full_fk_local_240k.robustness_e8_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 23797,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/robustness_e8_5.eval_robustness_e4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 18,
"total_frames": 10770,
"total_tasks": 1,
"total_videos": 36,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:18"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e4.robustness_e8This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 51,
"total_frames": 23580,
"total_tasks": 1,
"total_videos": 102,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/robustness_e8.path-vqa-robustnesseval_robustness_e9_3_full_dp_local_260kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 175,
"total_frames": 72626,
"total_tasks": 1,
"total_videos": 350,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:175"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e9_3_full_dp_local_260k.tokenization_robustness_v102
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/tokenization_robustness_v102.mechanistic-robustness
Decoding Robustness Results
Mechanistic robustness evaluation results for language models under six input
perturbations: character replacement, BPE-token replacement, word replacement,
local token shuffle, typographical corruption, and synonym replacement.
The repository is organized by model and perturbation:
models/<model>/<perturbation>/<percentage>/evals.csv
The qwen2.5_1.5b/adversarial directory contains the separate adversarial
evaluation outputs and manifest. Failed or… See the full description on the dataset page: https://huggingface.co/datasets/emizfliu/mechanistic-robustness.eval_robustness_e8_no2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 25,
"total_frames": 12633,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e8_no2.slake-robustnessprobe-robustness-rotten_tomatoes
rotten_tomatoes
Dataset repo: wrynx/probe-robustness-rotten_tomatoes
Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead.
Stats
Total records: 10662
Records per split:
test: 1066
train: 8530
valid: 1066
Number of classes: 2
Records per class:
0: 5331
1: 5331
Records per class per split:
test:
0: 533
1: 533
train:
0: 4265
1: 4265
valid:
0: 533
1: 533
Original dataset README… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-rotten_tomatoes.eval_robustness_e6_3_220kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 23551,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e6_3_220k.lighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230
Home Object Detection, Grasping and Sorting Eval — YOLOv8
Evaluation dataset for a robotic arm that detects, grasps, and sorts objects by type in home environments. 30 renders at 640x640 across kitchen, entry, living room and dressing spaces, staged with everyday household objects. Includes RGB plus metric depth, world-space normals (OpenGL, linear), albedo and material index passes, per-frame annotations, and midday lighting. Targets a YOLOv8 model.
This dataset mirrors public… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/lighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230.robustness_e3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 46,
"total_frames": 18597,
"total_tasks": 1,
"total_videos": 92,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:46"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/robustness_e3.eval_robustness_e8_no3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 25,
"total_frames": 12180,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e8_no3.openvla-oft-plus-robustness-recovery
OpenVLA-OFT+ Robustness and Recovery Dataset
成功した操作だけを増やせば、ロボット方策は頑健になるに違いない。
しかし、OpenVLA-OFT+の挙動を条件別に確かめると、失敗は一様ではなかった。
Spatialではsafe successが63.3%まで下がり、collision rateは35.0%に達した。
Goalのsafe success 90.0%、Objectの95.0%と比べても、空間関係を扱う操作の不安定さが際立つ。
このデータセットは、その差を学習データへ戻すために作成した。
遠距離の空間操作に加え、把持に失敗した状態、持ち上げられなかった状態、運搬中に物体を落とした状態から、成功まで立て直す軌跡を収録している。
弱点の観測
各suiteは20 scene、3条件、合計60エピソードで測定した。
95%信頼区間はscene単位のcluster bootstrap(10,000回)で算出している。
Suite
Episodes
Safe… See the full description on the dataset page: https://huggingface.co/datasets/argo11/openvla-oft-plus-robustness-recovery.robustness_e6_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 21,
"total_frames": 9757,
"total_tasks": 1,
"total_videos": 42,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/robustness_e6_2.eval_robustness_e8_no4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 25,
"total_frames": 12894,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/eval_robustness_e8_no4.
