datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oxford-iiit-pet
The Oxford-IIIT Pet Dataset
Description
A 37 category pet dataset with roughly 200 images for each class. The images have a large variations in scale, pose and lighting.
This instance of the dataset uses standard label ordering and includes the standard train/test splits. Trimaps and bbox are not included, but there is an image_id field that can be used to reference those annotations from official metadata.
Website: https://www.robots.ox.ac.uk/~vgg/data/pets/… See the full description on the dataset page: https://huggingface.co/datasets/timm/oxford-iiit-pet.redditstreet2shopworldsimprobe
WorldSimProbe Public Evaluation Inputs
This is the public input release for WorldSimProbe, an action-conditioned video
generation benchmark for embodied world models. It contains separate,
self-contained evaluation packages for RoboTwin, LIBERO, and ManiSkill.
Ground-truth videos, evaluator annotations, simulator outcomes, private ID maps,
and task-specific hidden labels are intentionally excluded. The official
evaluator holds those materials separately.
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/petersonco/worldsimprobe.krebs
The Krebs cycle dataset
Motivation
This dataset contains simulated time series that mimic Kreb's cycle.
The intent of the datasets is for causal discovery from multivariate
time series data, provide ground truth causal relationships as well as
allow testing on multiple scenarios, including many short time series,
few long time series, as well as relative data instead of absolute values.
The dataset was created at the Czech Technical University in Prague
as part of the… See the full description on the dataset page: https://huggingface.co/datasets/petrrysavy/krebs.nq_hotpotqa_trainimaginative-perception-token-pet-ipt
Citation
Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988):
@misc{bigverdi2026imaginativeperceptiontokensenhance,
title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models},
author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-pet-ipt.wds_vtab-petsoxford-pets
Oxford-IIIT Pet Dataset
Images from The Oxford-IIIT Pet Dataset. Only images and labels have been pushed, segmentation annotations were ignored.
Homepage: https://www.robots.ox.ac.uk/~vgg/data/pets/
License:
Same as the original dataset.
ViMed-PET-CT
ViMed-PET-CT
📅 Update: April 23, 2026
🐛 Bug Fixes: Corrected field mismatches (blank/missing fields) and date/filename inconsistencies. Restored missing metadata for patient 1701 (Dec 2023).
✨ New Feature: Added English translations of reports (/reports_en) using Gemma-4-26B-A4B-it.
ℹ️ About the dataset
🍴 Forked and optimized compression of dacthai2807/ViMed-PET, converting .npy and chunked zip files into .npz files.
📝 Better annotation and guideline.
📂… See the full description on the dataset page: https://huggingface.co/datasets/thainamhoang/ViMed-PET-CT.imaginative-perception-token-pet-eval-ai2thor
Citation
Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988):
@misc{bigverdi2026imaginativeperceptiontokensenhance,
title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models},
author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-pet-eval-ai2thor.oxford-iiit-pet-vl-enriched
Visualize on Visual Layer
Oxford-IIIT-Pets-VL-Enriched
An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues!
With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues help to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.imaginative-perception-token-pet-eval-habitat
Dataset Card for "habitat_perspective_eval"
More Information needed
Citation
Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988):
@misc{bigverdi2026imaginativeperceptiontokensenhance,
title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models},
author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-pet-eval-habitat.ViMed-PET-part1
Dataset description for three years: 2017, 2018, 2019
This dataset contains data from three years (2017, 2018, 2019). Each year has several month folders, which are named as THANG {month}.
Each year folder is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract all data folders.
Folder structure after extraction
Each folder named THANG {month} of a year is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2807/ViMed-PET-part1.SafeBuild-Bench
SafeBuild-Bench
arXiv:2608.00068 |
Paper (ACM DL) |
Project page |
Code
SafeBuild-Bench is a construction-safety benchmark for multimodal large language
models. It contains expert-verified construction-site images with task-specific
annotations for hazard identification and hazard description. It was published at
ACM SIGKDD 2026.
This Hugging Face package uses the standard imagefolder layout:
images/: benchmark images, one per metadata row… See the full description on the dataset page: https://huggingface.co/datasets/peter23333/SafeBuild-Bench.ViMed-PET-part2
Dataset description for year 2023
This dataset contains data from 8 months: January to September, except August, stored in the following folders respectively:
THANG 1
THANG 2
...
THANG 7
THANG 9
The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract the data folders.
Folder structure after extraction
Each folder named THANG {month} is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/Duc2305/ViMed-PET-part2.ViMed-PET-part3
Dataset description for year 2023
This dataset contains data from three months: October, November, and December, stored in the following folders respectively:
THANG 10
THANG 11
THANG 12
The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract the data folders.
Folder structure after extraction
Each folder named THANG {month} is divided into 3 subfolders, corresponding to 2… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2k/ViMed-PET-part3.CodeDance-RLlonghorizon-orchestrator-benchmark
LongHorizon Orchestrator Benchmark
An offline benchmark for the judgements a manipulation orchestrator delegates to a
vision-language model. An orchestrator wraps a frozen low-level policy and replaces a compound
instruction ("put everything in the bin") with a stream of single-object subtasks; to do so it must
plan (decompose the instruction into subtasks), verify (judge from pixels whether the
current subtask is finished), track state (know which goals are already done), and —… See the full description on the dataset page: https://huggingface.co/datasets/petkopetkov/longhorizon-orchestrator-benchmark.Nav-R1typescript-codeCUHK-PEDES
Dataset Card for "CUHK-PEDES"
More Information needed
blockgen-3d
BlockGen-3D Dataset
Overview
BlockGen-3D is a large-scale dataset of voxelized 3D models with accompanying text descriptions, specifically designed for text-to-3D generation tasks. By processing and voxelizing models from the Objaverse dataset, we have created a standardized representation that is particularly suitable for training 3D diffusion models.
Our dataset provides two types of representations: shape-only models represented as binary occupancy grids, and colored… See the full description on the dataset page: https://huggingface.co/datasets/PeterAM4/blockgen-3d.IndustryCorpus2_petrochemical
IndustryCorpus2: Petrochemicals
This repository contains the IndustryCorpus2: Petrochemicals domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year =… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_petrochemical.NavR1-30k
Dataset Card for "NavR1-30k"
More Information needed
youtube-commons-small
📺 YouTube-Commons-Small 📺
This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license.
Dataset Description
This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes.
Features
The dataset includes the following information for each video:
Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.dataclaw-peteromallet
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value
Sessions… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/dataclaw-peteromallet.korean-petitions
청와대 국민청원
데이터 출처: https://github.com/lovit/petitions_archive
크기: 651.8MB
sample
{
"category": "반려동물",
"begin": "2017-08-25",
"end": "2017-11-23",
"content": "길고양이들 밥주고있는 사람입니다. 최근에 동네주민과 트러블이 생겨 싸움이 일어났습니다. 길고양이들이 모여든다고 밥주지마라고 윽박지르셨습니다. 쓰레기봉투를 뜯는다거나 사람에게 해끼치거나 하지 않았습니다. 단순히 고양이가 모여드는게 싫답니다. 그럼 애들은 굶어죽어야하나요? 길고양이들이 맘놓고 쉬고 밥먹을 수 있는 환경이 전혀 없는데 무작정 밥안주고 물 안주면 얘네는 어떻게 하나요? 안그래도 수명도 짧은데다가 길고양이를 상대로 학대하는 사람들도 많은데 너무 가엾습니다. 강동구청은 고양이 급식소라고 만들어주셨던데 동네마다 한개씩이라도 만들어… See the full description on the dataset page: https://huggingface.co/datasets/heegyu/korean-petitions.my-dataclaw-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/my-dataclaw-data.gsm8k-chat
