datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
batik-processedSpatialQA-ESpatialQA-E is a robot manipulation dataset focusing on spatial relationship understanding.
Paper:
https://arxiv.org/abs/2406.13642
GitHub repo:
https://github.com/BAAI-DCAI/SpatialBot
SpatialBot-general QA, a VLM with precise depth understanding:
https://huggingface.co/RussRobin/SpatialBot
SpatialBench, the spatial understanding benchmark in general QA:
https://huggingface.co/datasets/RussRobin/SpatialBench
PhD-webdataset
PhD Webdataset
This repository contains the packaged version of PhD. For a detailed introduction to PhD, please visit the official website.
Overview
The PhD Webdataset is designed to facilitate easy access and usage of the PhD dataset. It includes various fields in 'json' key. The data in this repo is totally the same as in PhD.
Installation
Ensure you have Hugging Face's datasets library installed. You can install it via pip:
pip install datasets… See the full description on the dataset page: https://huggingface.co/datasets/AIMClab-RUC/PhD-webdataset.icm-data-temprukopys-curated-mvp-v2
RUKOPYS Curated MVP: Ukrainian Handwriting Recognition Dataset
RUKOPYS Curated MVP is a cleaned, task-ready derivative of
UkrainianCatholicUniversity/rukopys for Ukrainian handwritten
document AI. It turns the raw RUKOPYS release into reproducible artifacts for page-level
vision-language fine-tuning, crop-level transcription, and layout detection.
This dataset is designed for practical HTR work: train a model, inspect the normalized records,
evaluate layout/text extraction, and… See the full description on the dataset page: https://huggingface.co/datasets/AlexandreSheva/rukopys-curated-mvp-v2.rukopys-curated-mvp
RUKOPYS Curated MVP: Ukrainian Handwriting Recognition Dataset
Task-ready curated derivative of UkrainianCatholicUniversity/rukopys for Ukrainian handwritten document AI.
This dataset turns the raw RUKOPYS release into reproducible training artifacts for full-page vision-language fine-tuning, crop-level handwriting transcription, and layout detection. It was created to support an end-to-end HTR pipeline: curation, model training, inference, evaluation.
What This… See the full description on the dataset page: https://huggingface.co/datasets/AlexandreSheva/rukopys-curated-mvp.Conan-91kphysense_carla_dataset
PhySense CARLA Synthesized Dataset
Dataset Summary
This dataset is the CARLA-synthesized traffic dataset used in PhySense: Defending Physically Realizable Attacks for Autonomous Systems via Consistency Reasoning, CCS ’24. It is generated using the CARLA simulator and CARLA PythonAPI.
Source / Collection
The dataset is synthesized in CARLA. Instructions and scripts to reproduce/collect the dataset are provided in the accompanying GitHub repository:
Collection… See the full description on the dataset page: https://huggingface.co/datasets/Ruoyao/physense_carla_dataset.RUOK_ver_samsft_rule
