datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IMAGE_UNDERSTANDINGA key question for understanding multimodal performance is analyzing the ability for a model to have basic
vs. detailed understanding of images. These capabilities are needed for models to be used in
real-world tasks, such as an assistant in the physical world. While there are many dataset for object detection
and recognition, there are few that test spatial reasoning and other more targeted task such as visual prompting.
The datasets that do exist are static and publicly available, thus… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/IMAGE_UNDERSTANDING.kitchen-workspace-understanding-safe-manipulation
Kitchen Workspace Understanding & Safe Manipulation
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.multi-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.physical-ai-bench-understanding-evalsroie_document_understanding
Dataset Card for "sroie_document_understanding"
Dataset Description
This dataset is an enriched version of SROIE 2019 dataset with additional labels for line descriptions and line totals for OCR and layout understanding.
Dataset Structure
DatasetDict({
train: Dataset({
features: ['image', 'ocr'],
num_rows: 652
})
})
Data Fields
{
'image': PIL Image object,
'ocr': [
# text box 1
{
'box':… See the full description on the dataset page: https://huggingface.co/datasets/arvindrajan92/sroie_document_understanding.Spatial_Understanding
Purpose
Spatial intelligence is a fundamental component of both Artificial General Intelligence (AGI) and Embodied AI, encompassing multiple cognitive levels — Perception, Understanding, and Extrapolation (referring to the work).
We construct a composite benchmark derived from several prior works and this testbed is designed to measure the Understanding level of spatial intelligence of AI models within the given visual cues.
Overview
The benchmark integrates three… See the full description on the dataset page: https://huggingface.co/datasets/LLDDSS/Spatial_Understanding.humor_understanding_combinedPPTBench-Understanding
PPTBench Understanding Dataset
A collection of PowerPoint slides with associated understanding tasks and metadata.
Dataset Structure
The dataset contains the following fields for each entry:
hash: Unique identifier for each slide
category: Category/topic of the slide
task: Understanding task associated with the slide
description: Description of the slide content
question: Question formulated for the slide
Usage
Loading the Dataset
You can load this… See the full description on the dataset page: https://huggingface.co/datasets/tyrionhuu/PPTBench-Understanding.IMAGE_UNDERSTANDINGA key question for understanding multimodal performance is analyzing the ability for a model to have basic
vs. detailed understanding of images. These capabilities are needed for models to be used in
real-world tasks, such as an assistant in the physical world. While there are many dataset for object detection
and recognition, there are few that test spatial reasoning and other more targeted task such as visual prompting.
The datasets that do exist are static and publicly available, thus… See the full description on the dataset page: https://huggingface.co/datasets/neelsj/IMAGE_UNDERSTANDING.humor_understanding_nytmerged-expert-understanding-libero-datasetPPTBench-Understanding-testsroie-for-layout-understandinghumor_understanding_deepevalsarcasm_understanding_dpoGraphics2Code-Understanding
