datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
modified_libero_rlds
Modified LIBERO RLDS Datasets
This repository contains the four modified LIBERO datasets
used in the OpenVLA fine-tuning experiments, stored in RLDS data format. See Appendix E in the
OpenVLA paper for details about the fine-tuning experiments and
specific dataset modifications, and see the OpenVLA GitHub README
for instructions on how to run OpenVLA in LIBERO environments.
Citation
BibTeX:
@article{kim24openvla,
title={OpenVLA: An Open-Source… See the full description on the dataset page: https://huggingface.co/datasets/openvla/modified_libero_rlds.spreadsheet-bench-v2-modified
SpreadsheetBench V2 Modified: Multi-Document QA
1,060 questions and reference answers grounded in 127 Excel workbooks, 35 PDFs and 9 DOCX files. This independent derivative of SpreadsheetBench 2 shifts the task from editing spreadsheets and producing workbook deliverables toward finding, interpreting and combining information in business documents.
An independent project built entirely from publicly available source material and newly authored QA annotations. No private company… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/spreadsheet-bench-v2-modified.modified_libero_rlds
Modified LIBERO RLDS Datasets
This repository contains the four modified LIBERO datasets
used in the OpenVLA fine-tuning experiments, stored in RLDS data format. See Appendix E in the
OpenVLA paper for details about the fine-tuning experiments and
specific dataset modifications, and see the OpenVLA GitHub README
for instructions on how to run OpenVLA in LIBERO environments.
Citation
BibTeX:
@article{kim24openvla,
title={OpenVLA: An Open-Source… See the full description on the dataset page: https://huggingface.co/datasets/dachengzisks/modified_libero_rlds.pile-modified
Dataset Card for "pile-modified"
More Information needed
Crop-Yield-Prediction-MODIS
Crop Yield Prediction MODIS
This repository hosts the processed MODIS data used for crop-yield regression in the DFYP project. It was prepared from the MODIS branch of the DFYP project.
Repository: https://github.com/onef1shy/DFYP.
Paper: https://doi.org/10.1109/TGRS.2026.3684831
Contents
datasets/modis/processed_data/<year>/*.npy: preprocessed yearly samples indexed by year, county, and sample id
datasets/modis/processed_data/histogram_all_full.npz: histogram data… See the full description on the dataset page: https://huggingface.co/datasets/onef1shy/Crop-Yield-Prediction-MODIS.modiff-template-gallery
MoDiff Template Gallery
This public Dataset contains rights-approved media, input fixtures, posters,
and provenance records used by the open-source MoDiff template Gallery. MoDiff
executes local workflows with Hugging Face Diffusers and Modular Diffusers.
The Dataset exists so users can inspect or reuse the public examples without
adding large binary files to the application source repository.
Versioning and integrity
MoDiff releases pin this Dataset by its… See the full description on the dataset page: https://huggingface.co/datasets/sourav-das/modiff-template-gallery.Atmos_MODIS_AODmodified_libero_rlds
Modified LIBERO RLDS Datasets
This repository contains the four modified LIBERO datasets
used in the OpenVLA fine-tuning experiments, stored in RLDS data format. See Appendix E in the
OpenVLA paper for details about the fine-tuning experiments and
specific dataset modifications, and see the OpenVLA GitHub README
for instructions on how to run OpenVLA in LIBERO environments.
Citation
BibTeX:
@article{kim24openvla,
title={OpenVLA: An Open-Source… See the full description on the dataset page: https://huggingface.co/datasets/Chenyang5/modified_libero_rlds.modified_libero_rlds_cotdepsign_language_comparison_table_modifiedunitree-g1-hardware-modifications
Unitree G1 Hardware Modifications
Custom 3D-printable parts and reference photos for modifying the Unitree G1: OpenArm
grippers on the wrists, an extra head camera, and a backpack enclosure for the CAN-FD interface
and cabling.
Everything here is printable geometry plus photos of the assembled result — there is no
firmware or control code in this repo.
Modifications
1. Gripper on the G1 wrist
The gripper itself is the stock gripper from the OpenArm… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/unitree-g1-hardware-modifications.modified_libero_hdf5
Modified LIBERO HDF5 Data
This directory contains the modified version of LIBERO benchmark data. We downloaded the original LIBERO data and run the script from OpenVLA.
The modification is as follows:
image resolution from 128x128 to 256x256
filter out zero-valued actions (no-op) that do not change the robot's state
This modified data will be further processed by our preprocessing code.
modis-13q1-korea
MODIS 식생지수 16일 합성 자료: 대한민국 교차 타일
대한민국 육지 경계와 교차하는 modis-13Q1-061 데이터입니다.
합성: 16일 동안 촬영한 영상에서 픽셀마다 가장 적합한 관측값을 하나 선택
데이터 획득 방법
Microsoft Planetary Computer의 STAC Catalog로부터 COG 데이터와 XML 메타데이터를 다운로드했습니다. 별도로 clip하지 않았습니다.
STAC 검색 조건은 다음과 같습니다.
시간: 2000-01-01~2025-12-31
공간(타일): h27v05, h28v05 (대한민국과 교차하는 타일 선택)
Terra MOD13Q1과 Aqua MYD13Q1을 모두 포함
변수와 값 해석
[!NOTE]
아래 범위와 결측값은 scale factor 적용 전 정수값입니다.
결측값을 먼저 마스킹한 뒤 실제값 = 원시값 × scale을 적용해야 합니다.… See the full description on the dataset page: https://huggingface.co/datasets/husgbb/modis-13q1-korea.GAIA-modified
GAIA dataset
GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc).
We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format.
Data and leaderboard
GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/evatan/GAIA-modified.modified_libero_rlds
Modified LIBERO RLDS Datasets
This repository contains the four modified LIBERO datasets
used in the OpenVLA fine-tuning experiments, stored in RLDS data format. See Appendix E in the
OpenVLA paper for details about the fine-tuning experiments and
specific dataset modifications, and see the OpenVLA GitHub README
for instructions on how to run OpenVLA in LIBERO environments.
Citation
BibTeX:
@article{kim24openvla,
title={OpenVLA: An Open-Source… See the full description on the dataset page: https://huggingface.co/datasets/freshmanwang/modified_libero_rlds.modified_libero_rlds
Modified LIBERO RLDS Datasets
This repository contains the four modified LIBERO datasets
used in the OpenVLA fine-tuning experiments, stored in RLDS data format. See Appendix E in the
OpenVLA paper for details about the fine-tuning experiments and
specific dataset modifications, and see the OpenVLA GitHub README
for instructions on how to run OpenVLA in LIBERO environments.
Citation
BibTeX:
@article{kim24openvla,
title={OpenVLA: An Open-Source… See the full description on the dataset page: https://huggingface.co/datasets/wangyuanhao1213/modified_libero_rlds.openarms-hardware-modifications
OpenArm Hardware Modifications for Cloth Folding
Custom 3D-printable parts used in the Unfolding Robotics project, where we trained a bimanual robot to fold t-shirts with a 90% success rate.
These files modify the standard OpenArm.
Files
File
Description
J4_5cm_extended.step
Extended upper arm (bicep) segment, adds +5 cm of reach to compensate for the lack of a hip/torso in our setup. STEP format for easy modification.
J3-J4_Cover front extended.stl
Front… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/openarms-hardware-modifications.bugpilot-bugintro-lm-modify-gpt55-1k-oracleclean-443-20260506
BugPilot LM-Modify GPT-5.5 1k Oracle-Clean 443
This dataset contains the 443 task directories from the repaired LM-modify workspace that currently pass the oracle audit.
Source workspace: /data/augustine/demiurge/projects/experimental/training_swe_skrl_tinker/audits/lm_modify_target600_repair_workspace_20260506_v4
Source audit: full_oracle_postswaps_20260506_multinode32_c4
Export date: 2026-05-06
The directory layout matches the original task dataset layout: one task directory per… See the full description on the dataset page: https://huggingface.co/datasets/VmaxRL/bugpilot-bugintro-lm-modify-gpt55-1k-oracleclean-443-20260506.EmbSpatial-ModifiedVSI-Bench-modifiedmodified_libero_rlds
Modified LIBERO RLDS Datasets
This repository contains the four modified LIBERO datasets
used in the OpenVLA fine-tuning experiments, stored in RLDS data format. See Appendix E in the
OpenVLA paper for details about the fine-tuning experiments and
specific dataset modifications, and see the OpenVLA GitHub README
for instructions on how to run OpenVLA in LIBERO environments.
Citation
BibTeX:
@article{kim24openvla,
title={OpenVLA: An Open-Source… See the full description on the dataset page: https://huggingface.co/datasets/mww0531/modified_libero_rlds.ASR_french_modifiedPPTBench-Modificationduobench_modify
DuoBench Modify
duobench_modify is an unofficial derivative of
RobotControlStack/duobench.
It combines the 11 DuoBench simulation subsets into one LeRobot v3 dataset,
preserves the original joint-space fields, and adds 20-dimensional
end-effector (EEF) state and action fields for TwinVLA-style training.
Scope: this repository contains simulation data only. The four upstream
real-robot subsets are not included.
Dataset summary
Property
Value
Episodes… See the full description on the dataset page: https://huggingface.co/datasets/kisarakira/duobench_modify.modified-swiss-dwellings-enriched
Modified Swiss Dwellings (MSD), enriched
Floor plans of medium-to-large multi-apartment building complexes (ECCV 2024 benchmark
MSD), each linking three modalities of one floor plan:
image, geometry, and access graph.
1. Why this is here & what was done
Hosted on Hugging Face for reach and one-line loading by the ML community. The Swiss
Dwellings (SD) license (CC BY 4.0) permits redistribution with attribution — so this
also enriches the public MSD release, which… See the full description on the dataset page: https://huggingface.co/datasets/philippds/modified-swiss-dwellings-enriched.hotpot_qa_modifiedThis dataset is a modified version of the HotPotQA distractor dataset, which contains factual questions requiring multi-hop reasoning.
In the original HotPotQA dataset, each example presents ten paragraphs, only two of which contain the information necessary to answer the question; the remaining eight paragraphs include closely related but irrelevant details.
Consequently, solving this task requires the model to identify and reason over the pertinent passages.
To more strongly develop… See the full description on the dataset page: https://huggingface.co/datasets/markstanl/hotpot_qa_modified.Market1501-Background-Modified
Dataset Card for "Market1501-Background-Modified"
Dataset Summary
The Market1501-Background-Modified dataset is a variation of the original Market1501 dataset. It focuses on reducing the influence of background information by replacing the backgrounds in the images with solid colors, noise patterns, or other simplified alternatives. This dataset is designed for person re-identification (ReID) tasks, ensuring models learn person-specific features while ignoring background… See the full description on the dataset page: https://huggingface.co/datasets/ideepankarsharma2003/Market1501-Background-Modified.record-act-finetuned-base-modify-hold-posThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Tron-Hayato/record-act-finetuned-base-modify-hold-pos.sharc_modifiedShARC, a conversational QA task, requires a system to answer user questions based on rules expressed in natural language text. However, it is found that in the ShARC dataset there are multiple spurious patterns that could be exploited by neural models. SharcModified is a new dataset which reduces the patterns identified in the original dataset. To reduce the sensitivity of neural models, for each occurence of an instance conforming to any of the patterns, we automatically construct alternatives where we choose to either replace the current instance with an alternative instance which does not exhibit the pattern; or retain the original instance. The modified ShARC has two versions sharc-mod and history-shuffled. For morre details refer to Appendix A.3 .record-modify-hold-pos_20260910_100252This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Tron-Hayato/record-modify-hold-pos_20260910_100252.
