datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ImagePulseV2-Edit-Structure
ImagePulseV2 Dataset - Image Structure
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It comprises multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Structure.pseudo-camera-10k-structured-json
pseudo-camera-10k, structured JSON captions
The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled.
The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.bank-statement-structure-recognition
Synthetic Bank Statement Table Structure Dataset
A synthetically generated collection of bank statement images with pixel-perfect, automatically-produced bounding box annotations for table structure recognition (TSR).
🔑 In one sentence: fake bank statements + auto-generated YOLO labels for every table cell, built so you can train table-detection models (TATR, DETR, YOLO) without manual annotation.
At a Glance
Task
Object Detection → Table… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/bank-statement-structure-recognition.augmented-skin-images-50k-structuredstructured_imagesstructure_wildfire_damage_classification
Dataset Card for Structures Damaged by Wildfire
Homepage: Image Dataset of Structures Damaged by Wildfire in California 2020-2022
Dataset Summary
The dataset contains over 18,000 images of homes damaged by wildfire between 2020 and 2022 in California, USA, captured by the California Department of Forestry and Fire Protection (Cal Fire) during the damage assessment process. The dataset spans across more than 18 wildfire events, including the 2020 August Complex Fire, the… See the full description on the dataset page: https://huggingface.co/datasets/kevincluo/structure_wildfire_damage_classification.STARE-structured-analysis-of-the-retinaarXiv:2501.18921https://arxiv.org/abs/2501.18921
From_Reasoning_Structure_to_the_Ancient_Problem_of_Primes
From Reasoning Structure to the Ancient Problem of Primes
Author: Zixi Li (Oz Lee)
Date: 2025
Publisher: Hugging Face
Citation
@misc{oz_lee_2025,
author = { Oz Lee },
title = { From_Reasoning_Structure_to_the_Ancient_Problem_of_Primes (Revision d9034a1) },
year = 2025,
url = { https://huggingface.co/datasets/OzTianlu/From_Reasoning_Structure_to_the_Ancient_Problem_of_Primes },
doi = { 10.57967/hf/7156 }… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/From_Reasoning_Structure_to_the_Ancient_Problem_of_Primes.structured-vitalsbank-statement-structure-recognition
Synthetic Bank Statement Table Structure Dataset
A synthetically generated collection of bank statement images with pixel-perfect, automatically-produced bounding box annotations for table structure recognition (TSR).
🔑 In one sentence: fake bank statements + auto-generated YOLO labels for every table cell, built so you can train table-detection models (TATR, DETR, YOLO) without manual annotation.
At a Glance
Task
Object Detection → Table… See the full description on the dataset page: https://huggingface.co/datasets/sajid1235/bank-statement-structure-recognition.ocr-structure-datasetmetamdp-robosuite-franka-moving_ball-l3-structured-train-state16-h50-v1ANI1xmetamdp-robosuite-franka-moving_ball-l2-structured-train-state16-h50-v1song_structure_with_testsong_structure
Dataset Card for Song Structure
The raw dataset comprises 300 pop songs in .mp3 format, sourced from the NetEase music, accompanied by a structure annotation file for each song in .txt format. The annotator for music structure is a professional musician and teacher from the China Conservatory of Music. For the statistics of the dataset, there are 208 Chinese songs, 87 English songs, three Korean songs and two Japanese songs. The song structures are labeled as follows: intro… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/song_structure.structureqa数据集介绍
【图表,海报数据】chart_galaxy_full_4k.parquet 从ChartGalaxy/ChartGalaxy挑选出的图表 图标 海报的QA数据
【学术图数据】plotqa_sample_8k.parquet 从NiteshMethani/PlotQA: Dataset introduced in PlotQA: Reasoning over Scientific Plots中挑选出的学术图表数据 原始文件中包含100k 包含四个类别 现在的8k是四个类别随机抽样出2k 其中大量都是matplotlib等python代码生成的
【文档数据】tatdqa_full_11k.parquet 从NExTplusplus/TAT-DQA: TAT-DQA: Towards Complex Document Understanding By Discrete Reasoning中格式对齐后的文档QA 文档中包含表格 是更复杂的文档表格理解数据
【表格数据】the_cauldron_robut_wtq_38k.parquet… See the full description on the dataset page: https://huggingface.co/datasets/qixiangbupt/structureqa.wikipedia_structured_contentsnoisy_architectural_structures_all_over_the_worldSmolVLM_Essay_Structuredstructured_images_easyStructurEditBench_v3_filteredstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetamc-books-output-structure
Document OCR using dots.mocr
This dataset contains OCR results from images in markbaggett/amc-books using dots.mocr, a 3B multilingual model with SOTA document parsing and SVG generation.
Processing Details
Source Dataset: markbaggett/amc-books
Model: rednote-hilab/dots.mocr
Number of Samples: 303
Processing Time: 78.1 min
Processing Date: 2026-06-24 10:40 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch… See the full description on the dataset page: https://huggingface.co/datasets/markbaggett/amc-books-output-structure.structured-generation-information-extraction-vlms-openbmb-RLAIF-V-DatasetSquid_Game_Society_Structure_CN_EN.zipCopyright (c) 2025 蔡尔彬 (Ephraim Chuah)
本作品依据“署名-相同方式共享 4.0 国际(CC BY-SA 4.0)”授权发布。
您可以自由地:
✅ 共享 —— 在任何媒介以任何形式复制、转载本作品✅ 演绎 —— 修改、转换、在原作基础上创作新作品(包括用于AI训练)
只要您遵守以下许可条款:
📌 署名 —— 您必须给予适当的署名,提供指向本许可协议的链接,并说明是否对原始作品作了修改。您可以合理方式进行,但不得以任何方式暗示作者对您或您的使用作了背书。
📌 相同方式共享 —— 如果您再混合、转换或基于本作品进行创作,您必须采用与原始许可协议相同的许可条款发布您的贡献内容。
❌ 不得添加法律术语或技术措施限制他人做本协议允许的事情。
英文原文参考:
https://creativecommons.org/licenses/by-sa/4.0/
For any questions, please credit and contact:Ephraim Chuah (蔡尔彬) | happyboomplay@protonmail.com
cell_w2_structuredtable_structure_samplestructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetstructured3d-archWe provide parametric architectures for use with SSTK extracted from Structure3D.
|-- arch # JSON files specifying the [architecture](https://github.com/smartscenes/sstk/wiki/Architecture-Format)
|-- arch_renders # Renderings of the architecture
|-- roomId # Renderings colored by roomId
|-- images # png images
|-- camera_poses # Information about the camera paramters used for the rendering
|-- metadata… See the full description on the dataset page: https://huggingface.co/datasets/3dlg-hcvc/structured3d-arch.
