CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SQuADDS /SQuADDS_Layouts SQuADDS Layouts - versioned GDS artifacts for superconducting quantum hardware SQuADDS Layouts is the geometry-artifact companion to SQuADDS_DB, the Superconducting Qubit And Device Design and Simulation Database. It provides checksum-verified GDS files, stable geometry identities, and machine-readable geometry metadata so a simulation result can be traced to the exact layout that produced it. Homepage: https://lfl-lab.github.io/SQuADDS/ Repository:… See the full description on the dataset page: https://huggingface.co/datasets/SQuADDS/SQuADDS_Layouts.tabular10K<n<100K1 likes10k downloads24d agoHugging Face02weikaih /synthetic_data_v5_finegrain_layout_relight_with_our_synthetic_data_coco_l_full_500kimage1K<n<10K0 likes6.3k downloads2y agoHugging Face03nielsr /funsd-layoutlmv3imagen<1K42 likes1.3k downloads1y agoHugging Face04EditFigure /regulated_layout_dataset_v9_20260802text100K<n<1M0 likes876 downloads2mo agoHugging Face05nakamura196 /ndl-layout-dataset NDL-DocL Kotenseki Layout Dataset (YOLO format) A YOLO-formatted conversion of the kotenseki (pre-modern Japanese materials, 古典籍資料) subset of the NDL-DocL dataset published by the National Diet Library of Japan (NDL). Source dataset: https://github.com/ndl-lab/layout-dataset Source images: NDL Digital Collections https://dl.ndl.go.jp/ 国立国会図書館が公開する NDL-DocL データセットのうち、古典籍資料を YOLO 形式(Ultralytics 互換)に変換したものです。 This is a modified/derived version. The bounding boxes were converted… See the full description on the dataset page: https://huggingface.co/datasets/nakamura196/ndl-layout-dataset.imageobject-detection1K<n<10K1 likes821 downloads4d agoHugging Face06gzzyyxy /layout_diffusion_hypersimThis repository contains the data for SceneCraft: Layout-Guided 3D Scene Generation. Project page: https://orangesodahub.github.io/SceneCraft Code: https://github.com/OrangeSodahub/SceneCraft imagetext-to-3d10K<n<100K1 likes550 downloads1y agoHugging Face07HuiZhang0812 /LayoutSAM LayoutSAM Dataset Overview The LayoutSAM dataset is a large-scale layout dataset derived from the SAM dataset, containing 2.7 million image-text pairs and 10.7 million entities. Each entity is annotated with a spatial position (i.e., bounding box) and a textual description. Traditional layout datasets often exhibit a closed-set and coarse-grained nature, which may limit the model's ability to generate complex attributes such as color, shape, and texture.… See the full description on the dataset page: https://huggingface.co/datasets/HuiZhang0812/LayoutSAM.text1M<n<10M11 likes527 downloads2y agoHugging Face08OldDelorean /FloorplanQA-Layouts FloorplanQA (Layouts Only) This repository contains 2,000 JSON layouts used in the paper: "FloorplanQA: A Benchmark for Spatial Reasoning in LLMs Using Structured Representations" arXiv: https://arxiv.org/abs/2507.07644 Project Page: https://olddelorean.github.io/FloorplanQA/ Contents 600 synthetic kitchens 600 synthetic living rooms 600 synthetic bedrooms 200 layouts derived from HSSD-200 The release includes only the layouts. Structure layouts/… See the full description on the dataset page: https://huggingface.co/datasets/OldDelorean/FloorplanQA-Layouts.text1K<n<10K4 likes435 downloads4mo agoHugging Face09Extend-AI /RealDoc-Bench-Layout RealDocBench-Layout A 1,500-page document-layout benchmark for evaluating layout-detection models on real-world documents. COCO-style annotations across 9 block classes. Contents images/ — 1,500 page images (PNG / JPG / occasional WebP-as-PNG; see Caveats). annotations/<pageId>.json — per-page COCO files, each with a single image record, an annotations list, a categories list, and a page_info block. manifest.csv — pageId → domain + source URLs. The canonical row… See the full description on the dataset page: https://huggingface.co/datasets/Extend-AI/RealDoc-Bench-Layout.imageobject-detection1K<n<10K5 likes430 downloads4mo agoHugging Face10allenai /layout_distribution_shifttext10K<n<100K0 likes388 downloads3y agoHugging Face11Ssunbell /SORIE_layoutlmv2 Dataset Card for "SORIE_layoutlmv2" More Information needed text100K<n<1M0 likes382 downloads4y agoHugging Face12astrologos /docbank-layout Support the Project ☕ If you find this dataset helpful, please support me with a mocha: Dataset Summary DocBank is a large-scale dataset tailored for Document AI tasks, focusing on integrating textual and layout information. It comprises 500,000 document pages, divided into 400,000 for training, 50,000 for validation, and 50,000 for testing. The dataset is generated using a weak supervision approach, enabling efficient annotation of document structures… See the full description on the dataset page: https://huggingface.co/datasets/astrologos/docbank-layout.textgraph-ml100K<n<1M1 likes369 downloads2y agoHugging Face13R2aillc /LayoutOrderingHardimage1K<n<10K1 likes321 downloads2y agoHugging Face14surya-ai /question_LayoutLMimage0 likes286 downloads3y agoHugging Face15Spatial1ntelligence /synthetic-bedroom-layouts Synthetic Bedroom Layouts 4,000 procedurally composed bedroom layouts with 27,025 furniture placements, in the box format used by indoor scene-synthesis models (ATISS-style boxes.npz). Furniture is drawn from Amazon Berkeley Objects (CC BY 4.0), so the whole dataset is redistributable and usable commercially. Nothing here derives from 3D-FRONT, 3D-FUTURE, or any dataset that restricts redistribution. Where the numbers came from is written out, value by value, in PROVENANCE.md.… See the full description on the dataset page: https://huggingface.co/datasets/Spatial1ntelligence/synthetic-bedroom-layouts.3dother1K<n<10K3 likes282 downloads12d agoHugging Face16Sharka /DocVQA_LayoutLM_features Dataset Card for "DocVQA_LayoutLM_features" More Information needed 0 likes280 downloads3y agoHugging Face17j-min /layoutbench LayoutBench Release of LayoutBench dataset from Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation (CVPR 2024 Workshop) See also LayoutBench-COCO for zero-shot evaluation on OOD layouts with real objects. [Project Page] [Paper] Authors: Jaemin Cho, Linjie Li, Zhengyuan Yang, Zhe Gan, Lijuan Wang, Mohit Bansal Summary LayoutBench is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.… See the full description on the dataset page: https://huggingface.co/datasets/j-min/layoutbench.imagetext-to-image1K<n<10K1 likes269 downloads2y agoHugging Face18zhenyupan /3d_layout_reasoningDataset for MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse Github: https://github.com/PzySeere/MetaSpatial imageimage-text-to-textn<1K2 likes251 downloads1y agoHugging Face19EditFigure /regulated_layout_dataset_v10_20260830text100K<n<1M0 likes247 downloads14d agoHugging Face20SQuADDS /SQuADDS_Layout_Embeddings SQuADDS Layout Embeddings Versioned layout representations for the 24,106 GDS artifacts in SQuADDS/SQuADDS_Layouts. Static embedding model v0 static-embedding-v0 implements the original SQuADDS proof-of-concept model: v0 = parameter_sum + geometric_moments + flattened_shape_bitmap Each unit-normalized vector has 9,227 dimensions: Block Dimensions Contents Parameter sum 1 Permutation- and parameter-count-invariant sum of numerical design options… See the full description on the dataset page: https://huggingface.co/datasets/SQuADDS/SQuADDS_Layout_Embeddings.tabular10K<n<100K1 likes229 downloads24d agoHugging Face21Reza2kn /persian-ocr-community-dataset-layout Persian OCR Community Layout Annotations Resumable layout annotations for the page images in Reza2kn/persian-ocr-community-dataset. Each row points to an exact source dataset revision, Parquet shard, blob, and row. It includes the page identifier, page dimensions, handwriting flag, and structured layout boxes produced by datalab-to/surya_layout2 at confidence threshold 0.4. The boxes field contains label, confidence, raster-order position, and pixel coordinates x0, y0, x1, y1.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-community-dataset-layout.tabularobject-detection10K<n<100K0 likes217 downloads2mo agoHugging Face22runner21st /realm_layoutimage10K<n<100K0 likes202 downloads4mo agoHugging Face23KyleLin /LayoutPrompterA collection of datasets used in LayoutPrompter (NeurIPS2023). Specifically, publaynet and rico are downloaded from LayoutFormer++, posterlayout is downloaded from DS-GAN, and webui is downloaded from Parse-Then-Place. We sincerely thank them for the great work they do. image10K<n<100K3 likes198 downloads2y agoHugging Face24HuiZhang0812 /LayoutSAM-eval LayoutSAM-eval Benchmark Overview LayoutSAM-Eval is a comprehensive benchmark for evaluating the quality of Layout-to-Image (L2I) generation models. This benchmark assesses L2I generation quality from two perspectives: region-wise quality (spatial and attribute accuracy) and global-wise quality (visual quality and prompt following). It employs the VLM’s visual question answering to evaluate spatial and attribute adherence, and utilizes various metrics including IR score… See the full description on the dataset page: https://huggingface.co/datasets/HuiZhang0812/LayoutSAM-eval.image1K<n<10K2 likes190 downloads2y agoHugging Face25fimu-docproc-research /CIVQA-TesseractOCR-LayoutLM CIVQA TesseractOCR LayoutLM Dataset The Czech Invoice Visual Question Answering dataset was created with Tesseract OCR and encoded for the LayoutLM. The pre-encoded dataset can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR All invoices used in this dataset were obtained from public sources. Over these invoices, we were focusing on 15 different entities, which are crucial for processing the invoices. Invoice number Variable… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR-LayoutLM.2 likes184 downloads3y agoHugging Face26Sharka /DocVQA_layoutLM Dataset Card for "DocVQA_layoutLM" More Information needed tabular10K<n<100K0 likes179 downloads3y agoHugging Face27EditFigure /regulated_layout_dataset_v8_20260709text100K<n<1M0 likes165 downloads2mo agoHugging Face28bosungkim /layouts_procthor0 likes159 downloads11mo agoHugging Face29proustpunk /Layout_Synthetic_Datagatedimage1K<n<10K0 likes141 downloads11d agoHugging Face30magistermilitum /Tridis_layout_manuscripts A Unified Dataset for Codicological Document Layout Analysis Dataset Description This repository contains a large-scale, unified dataset for Document Layout Analysis (DLA) in historical manuscripts. It was created by harmonizing three distinct public corpora—e-NDP, CATMuS, and HORAE—which cover a wide range of document types from the 12th to the 17th century (administrative registers, literary manuscripts, printed books, and Books of Hours). The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/magistermilitum/Tridis_layout_manuscripts.image1K<n<10K0 likes140 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.