datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
terminal-bench-2
Terminal-Bench-2.0 Beta
Welcome to Terminal-Bench-2.0! If you’re reading this you’re a member of the Terminal-Bench community that we’ve selected to get a sneak peek at the latest version of the benchmark.
Getting Started
First, clone Harbor (formerly “Sandboxes”):
git clone https://github.com/laude-institute/harbor.git
From inside the Harbor directory run:
uv sync
This will install Harbor, our new package for running agent evals.
You should now be able to run TB 2.0!… See the full description on the dataset page: https://huggingface.co/datasets/penfever/terminal-bench-2.aloha_pen_uncap
Dataset Card for aloha_pen_uncap
This dataset is a FiftyOne conversion in LeRobot format of the aloha_pen_uncap_diverse subset of BiPlay.
The aloha_pen_uncap_diverse subset is a task-specific segment of BiPlay focusing on the long-horizon, dexterous bimanual task of un-capping a pen under diverse conditions. It contains episodes where the robot is required to grasp a pen and successfully remove its cap—an action requiring coordination and dexterity—across a wide range of object… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/aloha_pen_uncap.stable-diffusion-webui
Stable Diffusion web UI
A browser interface based on Gradio library for Stable Diffusion.
Features
Detailed feature showcase with images:
Original txt2img and img2img modes
One click install and run script (but you still must install python and git)
Outpainting
Inpainting
Color Sketch
Prompt Matrix
Stable Diffusion Upscale
Attention, specify parts of text that the model should pay more attention to
a man in a ((tuxedo)) - will pay more attention to tuxedo
a man in a… See the full description on the dataset page: https://huggingface.co/datasets/PennyJX/stable-diffusion-webui.PENCILablation_exploration_in_rl
Reinforcement Learning Improves Agentic Software Engineering
An ablation study of reinforcement-learning (RL) fine-tuning for agentic software-engineering (SWE) models. Starting from an 8B SFT model, we fine-tune with RL across ~20 configurations — varying the objective, loss normalization, sampling, and training dataset — and evaluate each on agentic SWE benchmarks.
Result
RL reliably and substantially improves agentic SWE performance, and the improvement is… See the full description on the dataset page: https://huggingface.co/datasets/penfever/ablation_exploration_in_rl.PENGWIN_Task2
PENGWIN Task 2: Pelvic Fragment Segmentation on Synthetic X-ray Images
Mirror of the training split of Task 2 of the MICCAI 2024 PENGWIN challenge
(https://pengwin.grand-challenge.org/), from the official Zenodo record
10913196 (train.zip, md5 9c90215dae54d8f494a85cfc7b19bc96).
These are SYNTHETIC X-rays, not real radiographs: DeepDRR renders of the 100 PENGWIN
Task 1 training CTs simulating intraoperative C-arm fluoroscopy, 500 random poses per CT
= 50,000 image/mask pairs.… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/PENGWIN_Task2.JANuS_datasetThis repository hosts the JANuS (Joint Annotations and Names) dataset introduced in the 2023 paper Distributionally Robust Classification on a Data Budget.
As of this writing, ours is the only public dataset which is both fully annotated with ground-truth labels and fully captioned with web-scraped captions.
It is designed to be used for controlled experiments with vision-language models.
What is in JANuS?
JANuS provides metadata and image links for four new training datasets; all… See the full description on the dataset page: https://huggingface.co/datasets/penfever/JANuS_dataset.aloha_pen_uncap_diverseThis dataset was created using LeRobot.
Dataset Description
This dataset is a lerobot conversion of the aloha_pen_uncap_diverse subset of BiPlay.
BiPlay contains 9.7 hours of bimanual data collected with an aloha robot at the RAIL lab @ UC Berkeley, USA. It contains 7023 clips, 2000 language annotations and 326 unique scenes.
Paper: https://huggingface.co/papers/2410.10088 Code: https://github.com/sudeepdasari/dit-policy If you use the dataset please cite:… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/aloha_pen_uncap_diverse.PENGWIN_Task1
PENGWIN Task 1 — Pelvic Fracture Segmentation on CT
The CT task of the PENGWIN 2024 challenge (PElvic bone fraGment (WIN)dow,
MICCAI 2024): segment the sacrum, left hipbone and right hipbone, and the
individual fracture fragments of each, in preoperative pelvic trauma CT.
This is an instance segmentation task, not a 3-class semantic one — the
label value identifies which fragment of which bone, and the fragment count
varies per case.
What this mirror contains — read… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/PENGWIN_Task1.Penguin-Recap-I
Penguin-Recap-I
Penguin-Recap-I publishes recap metadata only. The repository does not contain
image binaries.
Included subsets
subset
collection
local source roots
expected records
datacomp_coyo_penguin
DataComp + COYO Penguin recap
datamultimodal/IMAGE/datacomp_1b, datamultimodal/IMAGE/coyo_700m
57,618,155
sa1b_penguin
SA-1B Penguin recap
datamultimodal/IMAGE/SA-1B
9,254,501
openimages_penguin
OpenImages Penguin recap… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-I.BackMarble-5hand-gesture
Dataset Card for "hand-gesture"
More Information needed
pen_uncapThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 80,
"total_frames": 51100,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vo2yager/pen_uncap.pen-datasetgrip_pen_0711transportgs-blender-scenes
TransportGS Blender Scenes
One continuous RGB input stream at 2× speed, with no cut (sequence_000).
These synthetic RGB-D sequences study streaming Gaussian reconstruction when an object moves outside the camera's view. The camera first observes the object, turns away during the move, then returns for a partial revisit. A method must use the observation history to update the object at its new location and clear its old location. no_change sequences test whether it avoids… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/transportgs-blender-scenes.stationery_dataset_Eraser_scale_pencilsim-pen
ShadowHandPen RGB + raw rigid-contact tactile trajectories
This dataset contains 5,000 successful ShadowHandPen PPO rollouts. Each
episode stores synchronized RGB frames, robot hand joint/object state in
trajectory_env0.npz, and dense EgoTouch-layout pressure in
pressure_grids.npz. The pressure grids are computed from pairwise rigid
contacts and Gaussian projection onto 217 valid taxels per hand.
The tar files preserve the original collection shards. Shard-level duplicate
RGB… See the full description on the dataset page: https://huggingface.co/datasets/qqyang/sim-pen.multimath-300kuncap_the_red_pen_roboreward_retrievalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "generic",
"total_episodes": 100,
"total_frames": 12506,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30.0,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ykorkmaz/uncap_the_red_pen_roboreward_retrieval.Piper_Uncap_Penpiper_uncap_penThis dataset was created using LeRobot.
About Task
Task Objective: Pick up the pen from the desk and uncap the pen.
Operational Objects: Marker Pen
Operation Duration: Each operation takes approximately 15 to 20 seconds.
Recording Frequency: 15 Hz.
Robot Type: 7-DOF dual-arm Agilex 2 Pipers Desktop Robot.
End Effector: Gripper.
Dual-Arm Operation: Yes.
Image Resolution: 640x480.
Camera Positions: High; Low; Left; Right.
Data Content: • Robot's current state. •… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/piper_uncap_pen.counterfactual-pendulum-multilingual
📌 Dataset Summary
When a Vision-Language Model (VLM) is given an image along with a text prompt containing contradictory or misleading information, how does it react? Does it rely on the visual evidence, succumb to textual bias, or honestly abstain when faced with unresolvable conflict?
This dataset adapts the Counterfactual Pendulum scenario across two visual conflict dimensions:
Angular (Angle): Conflict in the pendulum's angle of inclination.
Light: Conflict in the light… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/counterfactual-pendulum-multilingual.so101_color_pen_sortverbosity-materials
Verbosity Research Materials
公开研究材料库:7篇论文、8个任务(4个website、4个slides)、40份作品(8份作者作品、32份AI作品)。
开始阅读 / Start here
下载完整 v1 ZIP(约288 MiB)
浏览展开目录
中文入门及离线限制
作品与模型来源清单
论文及作者原始链接
解压后打开 verbosity-materials-v1/index.html 浏览。HF文件仓库不作为这些HTML的运行网站。
Slides包含20份PDF,AI作品另附可编辑HTML和本地资源;作者slides仅有PDF。
论文正文是生成时使用的快照,完整论文通过原始链接访问。
Provenance and use
The v1 snapshot comes from run materials-share-20260918a. All 40 artifacts retain
provenance; the bundle includes SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/Penguin-N/verbosity-materials.realman_chargegun_30_penalty_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 30,
"total_frames": 1756,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yfynb1111/realman_chargegun_30_penalty_v2.uncap_the_red_pen_preference_retrievalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "generic",
"total_episodes": 151,
"total_frames": 16635,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:151"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/ykorkmaz/uncap_the_red_pen_preference_retrieval.pendulum-conflict
Pendulum Conflict Dataset
This dataset is a manually curated subset of 100 isolated multimodal conflict examples derived from the akomand/counterfactual_pendulum dataset.
It is designed for evaluating Multimodal Large Language Models (MLLMs) under controlled visual-textual conflicts.
Dataset Statistics
Total Samples: 100
Categories: angle (25), light (25), shadow_len (25), shadow_pos (25)
Language: English
Corrected / Audited Samples: 20
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/pendulum-conflict.uncap_the_red_pen_similarity_retrievalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "generic",
"total_episodes": 150,
"total_frames": 17049,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:150"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/ykorkmaz/uncap_the_red_pen_similarity_retrieval.
