datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
raw_primitive_datasets
Datacard
This is the official fine-tuning dataset provided by VLABench (raw data), with 500 episodes each task. The current version includes 10 primitive tasks.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
Download all archive files and use the following command to extract:
cat vlabench_primitive.tar.gz.* | tar -xzvf -
In the resulting… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/raw_primitive_datasets.IXI-DatasetsOpenS2S_Datasets
How to Use?
Download, merge the files, and extract
You can run the following command to merge the compressed file parts after downloading.
cat en_response_wav.tar.gz.* > en_response_wav.tar.gz
cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz
FRoM-W1-Datasets
FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions
The Humanoid Intelligence Team from FudanNLP and OpenMOSS
Introduction
Humanoid robots are capable of performing various actions such as greeting, dancing and even backflipping. However, these motions are often hard-coded or specifically trained, which limits their versatility. In this work, we present FRoM-W1[^1]… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FRoM-W1-Datasets.vigorl_datasets
ViGoRL Datasets
This repository contains the official datasets associated with the paper "Grounded Reinforcement Learning for Visual Reasoning (ViGoRL)", by Gabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain, Michael J. Tarr, Aviral Kumar, and Katerina Fragkiadaki.
Dataset Overview
These datasets are designed for training and evaluating visually grounded vision-language models (VLMs).
Datasets are organized by the visual reasoning tasks described in the ViGoRL… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/vigorl_datasets.sam2act-datasets
SAM2Act
SAM2Act is a multi-view robotics transformer policy for robotic manipulation. Built on RVT-2, it combines multi-resolution upsampling with visual embeddings from the SAM2 foundation model to improve 3D action prediction, multitask learning, and generalization. SAM2Act+ extends this policy with a memory bank, memory encoder, and memory attention so the agent can condition on prior observations and actions for spatial memory-dependent tasks.
For full project details, code… See the full description on the dataset page: https://huggingface.co/datasets/hqfang/sam2act-datasets.AnyDepth_datasetsVRF-datasetsGATE-VLAP-datasets
GATE-VLAP Datasets
Grounded Action Trajectory Embeddings with Vision-Language Action Planning
This repository contains preprocessed datasets from the LIBERO benchmark suite in WebDataset TAR format, specifically designed for training vision-language-action models with semantic action segmentation.
Data Format: WebDataset TAR
We provide datasets in WebDataset TAR format for optimal performance:
✅ Fast loading - Efficient streaming during training✅ Easy downloading - Single… See the full description on the dataset page: https://huggingface.co/datasets/gate-institute/GATE-VLAP-datasets.datasetsWan_datasets
rCM: Score-Regularized Continuous-Time Consistency Model
Paper | Website | Code
This repo holds Wan-synthesized datasets used for rCM training.
Citation
@article{zheng2025rcm,
title={Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency},
author={Zheng, Kaiwen and Wang, Yuji and Ma, Qianli and Chen, Huayu and Zhang, Jintao and Balaji, Yogesh and Chen, Jianfei and Liu, Ming-Yu and Zhu, Jun and Zhang, Qinsheng},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/hackroot/Wan_datasets.nanostep-datasetsGeoZero_Train_Datasets
SFT and RL Traning dataset of GeoZero
Dataset Composition
GeoZero consists of three variants:
File
Description
GeoZero-Raw.json
Raw aggregated data across heterogeneous datasets
GeoZero-Instruct.json
Unified instruction-tuned dataset for supervised fine-tuning
GeoZero-Hard.json
Challenging subset for RL training
All image files are stored under the images/ directory.
Directory Structure
GeoZero_Train_Datasets/
├── images/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/hjvsl/GeoZero_Train_Datasets.CFNet_Datasetsbalanced-audio-score-datasets-DACVAEGeoZero_Train_Datasets
SFT and RL Traning dataset of GeoZero
Dataset Composition
GeoZero consists of three variants:
File
Description
GeoZero-Raw.json
Raw aggregated data across heterogeneous datasets
GeoZero-Instruct.json
Unified instruction-tuned dataset for supervised fine-tuning
GeoZero-Hard.json
Challenging subset for RL training
All image files are stored under the images/ directory.
Directory Structure
GeoZero_Train_Datasets/
├── images/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/keshavagarwal2004dev/GeoZero_Train_Datasets.labelgs_datasetspatchman_datasetsvigorl_gaze_datasetslegal_datasets
Legal Case Documents Dataset Snapshot for GaiaNet
This repository contains a Qdrant database snapshot of vectorized legal case documents from Nigeria and the UK. The data has been processed and converted into embeddings suitable for use as a knowledge base in a Retrieval-Augmented Generation (RAG) system.
Purpose
This dataset was created to serve as a specialized knowledge base for the GaiaNet ecosystem. The goal is to enable a Large Language Model (LLM) to answer… See the full description on the dataset page: https://huggingface.co/datasets/Tobivictor/legal_datasets.pwm_datasetsdoc-audio-11
[doc] audio dataset 11
This dataset contains two tar files that contain pairs of samples with one audio file and one JSON file.
Catena_datasetspatchman_datasetsimg-demoSOC-defection-datasetsmy-gec-datasetsdatasets_49NL-To1_sft_datasetsdataset_split
