datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.ur5fail_test_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.bdv2fail_test_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.TestDatasetScratch repo for testing dataset-viewer schema inference. All data is synthetic and
contains no real personal information.
case_bank/ holds a few raw per-case JSON files whose nested mock block is union-typed
across tools (some lists are empty, some hold structs), which breaks single-schema
inference. viewer/cases.jsonl is a flattened table with a stable schema, and the
configs block above points the viewer at it so the raw files are not globbed.
testdatatest_data_202608010526003995_staging3_0507_test_dataset_final
3_0507_test_dataset_final
This dataset contains medical test questions and answers.
Files
data.json: Array of objects with keys like question_id, question, choices, correct_answer, etc.
test_dataEcom-Chatbot-Synthetic-Test-Datasettest-datasetschweinhunds_test_datasetscored_test_data_by_rm_filterroleplay-test
roleplay-test
Test split for lm-eval-harness.
MBPP-cluster_0-based-fewshot-prompting-test-dataset-take2test-datasetdata-50k-test-finaldata-50k-refine-test-gen2testdataset2MBPP-cluster-based-fewshot-prompting-test-datasettsp_testdataEnglish_Skills_Test_DatasetThis is our testing dataset for the skills. It contains all skill groups of the skills in our query samples.
TestDatasetschema_test_data_202602191021032182
