datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oneformer_demobridge-rlds
Dataset Structure
These datasets are used for MemoryVLA training.
This is the standard setting and can be directly used for other models as well.All data follow the RLDS format from the Bridge dataset.
bridge_orig — 60k+ episodes, widowx robot
ship-tracking-databehavior-1k_2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dario-Shit4/behavior-1k_2025-challenge-demos.alpaca-zh
Dataset Card for "alpaca-zh"
本数据集是参考Alpaca方法基于GPT4得到的self-instruct数据,约5万条。
Dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
It is the chinese dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM/blob/main/data/alpaca_gpt4_data_zh.json
Usage and License Notices
The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/alpaca-zh.shiur-clips-flactweet_temporal_shift"""
_TWEET_TEMPORAL_CITATION =Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.say-idc-media-vaultViT-FineTuneSTUZero-Atari-Dynamics
STUZero Atari Dynamics Dataset
Offline dynamics training datasets collected from trained EfficientZero V2 (EZv2) benchmark models on Atari games. Each game's data is stored in a subfolder named {game}_{steps} indicating the game and the number of training steps of the source checkpoint. While all models were trained for 120K steps, best results in some games were attained at earlier checkpoints. The model with best eval scores was used to curate data for each game.… See the full description on the dataset page: https://huggingface.co/datasets/Shivamkak/STUZero-Atari-Dynamics.cleanvid-15m_map
CleanVid Map (15M) 🎥
TempoFunk Video Generation Project
CleanVid-15M is a large-scale dataset of videos with multiple metadata entries such as:
Textual Descriptions 📃
Recording Equipment 📹
Categories 🔠
Framerate 🎞️
Aspect Ratio 📺
CleanVid aim is to improve the quality of WebVid-10M dataset by adding more data and cleaning the dataset by dewatermarking the videos in it.
This dataset includes only the map with the urls and metadata, with 3,694,510 more entries than… See the full description on the dataset page: https://huggingface.co/datasets/shinonomelab/cleanvid-15m_map.openloris-scene
OpenLORIS-Scene Datasets
This page is an index of the OpenLORIS-Scene datasets.
Terms of Use
The OpenLORIS-Scene datasets are released with the CC BY-ND 4.0 license, which means you can do anything with the data, even for commercial purposes, except distributing derivative datasets (contact us at openloris@gmail.com if you would like to do so). We would appreciate it if you citeour paper when appropriate.
Cite
X Shi, D Li et al. “Are We Ready for Service… See the full description on the dataset page: https://huggingface.co/datasets/shixuesong/openloris-scene.ZDPShift
ZDPShift: Beyond the Zero-Disparity Plane in Stereo
Every public stereo benchmark assumes positive disparity valuesd = fB/Z ≥ 0. Mordern stereoscopic display — cinema 3D, VR, HMDs — actively uses d < 0. ZDPShift bridges the gap: the same artist-authored open-movie content rendered at five ZDP shifts Δ ∈ {−16, 0, +16, +24, +32} pixels, giving you a controlled continuum from textbook-positive to substantially-crossed disparities, with analytical ground truth at every pixel.… See the full description on the dataset page: https://huggingface.co/datasets/shijianjian/ZDPShift.Challenge-phase1-dataset-rlinflibero-rlds
Dataset Structure
These datasets are used for MemoryVLA training.
This is the standard LIBERO setting and can be directly used for other models as well.All data follow the RLDS format from the LIBERO benchmark, where each task initially contains 50 trajectories and failed rollouts are filtered out.NOTE: LIBERO-90 is also included.
libero_spatial_no_noops — 10 tasks
libero_object_no_noops — 10 tasks
libero_goal_no_noops — 10 tasks
libero_10_no_noops — 10 tasks… See the full description on the dataset page: https://huggingface.co/datasets/shihao1895/libero-rlds.vae_cache_minecraft_480p_9sentity-explanationphysical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.shiunjikenokodomotachi
Bangumi Image Base of Shiunji-ke No Kodomotachi
This is the image base of bangumi Shiunji-ke no Kodomotachi, we detected 39 characters, 4423 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shiunjikenokodomotachi.sharegpt_gpt4
Dataset Card
Dataset Summary
ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。
Languages
数据集是多语言,包括中文、英文、日文等常用语言。
Dataset Structure
Data Fields
The data fields are the same among all splits.
conversations: a List of string .
head -n 1 sharegpt_gpt4.jsonl
{"conversations":[
{'from': 'human',
'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.gpt-4v-distribution-shift
License
This repository is licensed under the MIT License.
Description
This Hugging Face repository hosts the random case dataset utilized in our research project, detailed in the GitHub repository gpt-4v-distribution-shift.
These datasets are crucial for evaluating the performance of multimodal foundation models under various distribution shift scenarios.
Using the Dataset
For detailed instructions on how to use this dataset to reproduce the results presented… See the full description on the dataset page: https://huggingface.co/datasets/jameszhou-gl/gpt-4v-distribution-shift.robotwin_scan_object_place_dual_shoes_serialized_200medical纯文本数据,中文医疗数据集,包含预训练数据的百科数据,指令微调数据和奖励模型数据。ISSAI_KSC_335RS_v_1_1
Dataset Card for "ISSAI_KSC_335RS_v_1_1"
Kazakh Speech Corpus (KSC)
Identifier: SLR102
Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours)
Category: Speech
License: Attribution 4.0 International (CC BY 4.0)
Downloads (use a mirror closer to you):
ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN]
About this resource:
A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.shinka-cvdp-benchmark-fullbge-m3-data
Dataset Summary
This depository contains all the fine-tuning data for the bge-m3 model, including:
Dataset
Language
MS MARCO
English
NQ
English
HotpotQA
English
TriviaQA
English
SQuAD
English
COLIEE
English
PubMedQA
English
NLI from SimCSE
English
DuReader
Chinese
mMARCO-zh
Chinese
T2Ranking
Chinese
Law-GPT
Chinese
cMedQAv2
Chinese
NLI-zh
Chinese
LeCaRDv2
Chinese
Mr.TyDi
11 languages
MIRACL
16 languages
MLDR
13 languages
Note: The… See the full description on the dataset page: https://huggingface.co/datasets/Shitao/bge-m3-data.Shionone84MLDR
Dataset Summary
MLDR is a Multilingual Long-Document Retrieval dataset built on Wikipeida, Wudao and mC4, covering 13 typologically diverse languages. Specifically, we sample lengthy articles from Wikipedia, Wudao and mC4 datasets and randomly choose paragraphs from them. Then we use GPT-3.5 to generate questions based on these paragraphs. The generated question and the sampled article constitute a new text pair to the dataset. The prompt for GPT3.5 is “You are a curious AI… See the full description on the dataset page: https://huggingface.co/datasets/Shitao/MLDR.RGB-Event-ISP-Dataset
