CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01weikaih /synthetic_data_v5_finegrain_layout_relight_with_our_synthetic_data_coco_l_full_500kimage1K<n<10K0 likes6.4k downloads2y agoHugging Face02EditFigure /regulated_layout_dataset_v9_20260802text100K<n<1M0 likes876 downloads2mo agoHugging Face03nakamura196 /ndl-layout-dataset NDL-DocL Kotenseki Layout Dataset (YOLO format) A YOLO-formatted conversion of the kotenseki (pre-modern Japanese materials, 古典籍資料) subset of the NDL-DocL dataset published by the National Diet Library of Japan (NDL). Source dataset: https://github.com/ndl-lab/layout-dataset Source images: NDL Digital Collections https://dl.ndl.go.jp/ 国立国会図書館が公開する NDL-DocL データセットのうち、古典籍資料を YOLO 形式(Ultralytics 互換)に変換したものです。 This is a modified/derived version. The bounding boxes were converted… See the full description on the dataset page: https://huggingface.co/datasets/nakamura196/ndl-layout-dataset.imageobject-detection1K<n<10K1 likes821 downloads5d agoHugging Face04EditFigure /regulated_layout_dataset_v10_20260830text100K<n<1M0 likes270 downloads15d agoHugging Face05Reza2kn /persian-ocr-community-dataset-layout Persian OCR Community Layout Annotations Resumable layout annotations for the page images in Reza2kn/persian-ocr-community-dataset. Each row points to an exact source dataset revision, Parquet shard, blob, and row. It includes the page identifier, page dimensions, handwriting flag, and structured layout boxes produced by datalab-to/surya_layout2 at confidence threshold 0.4. The boxes field contains label, confidence, raster-order position, and pixel coordinates x0, y0, x1, y1.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-community-dataset-layout.tabularobject-detection10K<n<100K0 likes224 downloads2mo agoHugging Face06EditFigure /regulated_layout_dataset_v8_20260709text100K<n<1M0 likes162 downloads2mo agoHugging Face07proustpunk /Layout_Synthetic_Datagatedimage1K<n<10K0 likes142 downloads11d agoHugging Face08BowenC /Orchestra_Layout_Dataset Orchestral Score Layout Dataset Ground-truth annotations for staff layout detection in orchestral music scores, covering two Tchaikovsky symphonies. Contents dataset_layout/ ├── Tchai_4.pdf # Source PDF — Tchaikovsky Symphony No. 4 ├── Tchai_4/ # Page images (PNG, one per page) ├── Tchai_4_csv_gt/ # Per-page CSV ground truth for Tchai_4 │ ├── Tchai_6.pdf # Source PDF — Tchaikovsky Symphony No. 6 ├── Tchai_6/… See the full description on the dataset page: https://huggingface.co/datasets/BowenC/Orchestra_Layout_Dataset.imagen<1K0 likes132 downloads4mo agoHugging Face09raphael0202 /ingredient-detection-layout-dataset Dataset Card for "ingredient-detection-layout-dataset" More Information needed image1K<n<10K0 likes130 downloads3y agoHugging Face10umesh16071973 /Layout_Dataimagen<1K0 likes114 downloads3y agoHugging Face11Reza2kn /persian-ocr-community-dataset-layout-mergedtext1K<n<10K0 likes103 downloads2mo agoHugging Face12jjobear /collage-layout-dataset Collage Layout Synthetic Dataset Synthetic photo-collage layouts for layout-quality analysis & correction, built on a six-ingredient framework (Format, Photos, Visual Weight, Hierarchy, Readability, Harmony). Corrector-not-generator: every collage carries a naive (v1_center) and a corrected (fit) placement, so a model can learn the correction. Faces are synthetically replaced (privacy-safe). How to load from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/jjobear/collage-layout-dataset.imageimage-to-imagen<1K0 likes53 downloads2mo agoHugging Face13medieval-data /mgh-critical-edition-layout MGH Layout Detection Dataset Dataset Description General Description This dataset consists of scans from the MGH critical edition of Alcuin's letters, which were first edited by Ernestus Duemmler in 1895. The digital scans were sourced from the DMGH's repository, which can be accessed here. The scans were annotated using CVAT, marking out two classes: the title of a letter and the body of the letter. Why was this dataset created? The primary… See the full description on the dataset page: https://huggingface.co/datasets/medieval-data/mgh-critical-edition-layout.imagen<1K0 likes47 downloads3y agoHugging Face14jonny122 /khmer-newspaper-layout-dataset Khmer Newspaper Layout Dataset Dataset Description This dataset contains Khmer newspaper layouts with annotated regions for document layout analysis and OCR tasks. Dataset Summary Total Examples: 9,344 newspaper layouts Language: Khmer (Cambodian) Image Format: PNG Annotations: LabelMe JSON format with bounding boxes and segmentation masks Features config_id: Unique identifier for each sample image: Newspaper layout image (PNG)… See the full description on the dataset page: https://huggingface.co/datasets/jonny122/khmer-newspaper-layout-dataset.imageobject-detection1K<n<10K1 likes46 downloads7mo agoHugging Face15rinabuoy /layout-data-khmer-3gatedimage1K<n<10K0 likes41 downloads1mo agoHugging Face16Nunatic /dream-layout-svg-dataset-v11016 images of Random Boxs Along Bézier Curves image1K<n<10K2 likes38 downloads2y agoHugging Face17manu /illuin_layout_dataset_text_only Dataset Card for "illuin_layout_dataset_text_only" More Information needed text100K<n<1M0 likes37 downloads3y agoHugging Face18muthuk1 /alwas-analog-layout-dataset ALWAS Analog Layout Dataset Synthetic dataset for training ML models in the ALWAS (Analog Layout Workflow Automation System) pipeline. Dataset Description 4,000 analog IC layout blocks with complete metadata, stage transitions, and labels for: Hours estimation — actual vs estimated hours Complexity classification — Low / Medium / High Bottleneck risk prediction — Low / Medium / High Completion time prediction — stage-by-stage transition history Dataset… See the full description on the dataset page: https://huggingface.co/datasets/muthuk1/alwas-analog-layout-dataset.tabular1K<n<10K1 likes34 downloads5mo agoHugging Face19vinicios94 /elementor-layout-vlm-dataset Elementor Layout VLM Dataset 📊 Dataset Summary Dataset para fine-tuning de modelos Vision-Language (VLM) para geração de layouts Elementor a partir de imagens. Task: Visual Question Answering (VQA) Format: VLM VQA (image, question, answer) Total: 30 exemplos Training: 24 exemplos Validation: 6 exemplos 🎯 Uso Recomendado AutoTrain Configuration Task: VLM VQA Base Model: google/paligemma-3b-pt-448 Dataset: vinicios94/elementor-layout-vlm-dataset… See the full description on the dataset page: https://huggingface.co/datasets/vinicios94/elementor-layout-vlm-dataset.imagevisual-question-answeringn<1K0 likes32 downloads1y agoHugging Face20EditFigure /regulated_layout_dataset_v7_20260705image10K<n<100K0 likes26 downloads3mo agoHugging Face21milord-x /KazNU-OCR-Layout-dataset KazNU-OCR-Layout Dataset Датасет для layout detection (определение структурных блоков) страниц казахской университетской газеты «Qazaq Universiteti». Используется для обучения модели KazNU-OCR-Layout в рамках hybrid active learning пайплайна OCR газетного архива. 65 страниц, отрендеренных из оригинальных PDF выпусков (300 DPI), с ручной разметкой структурных блоков и их границ (bounding box). Состав images/ — PNG-страницы (300 DPI) annotations/ — JSON-файлы… See the full description on the dataset page: https://huggingface.co/datasets/milord-x/KazNU-OCR-Layout-dataset.object-detection0 likes15 downloads3mo agoHugging Face22Reza2kn /persian-ocr-community-dataset-layout-argilla persian-ocr-community-dataset-layout-argilla This dataset contains only two columns for Argilla import: image and label. Total pages (rows): 5878 Bboxes with OCR text: 33486 Bboxes without OCR text: 0 Pages with at least one OCR: 5871 Pages with zero OCR: 7 text1K<n<10K0 likes14 downloads2mo agoHugging Face23Nunatic /dream-layout-svg-dataset-v5 random rotation random translation even length sampling along spline curve (instead of using t directly) no-overlap image0 likes13 downloads2y agoHugging Face24saisriteja /LayoutData0 likes13 downloads6mo agoHugging Face25LIAGM /Bi_Layout_Datasetimage0 likes12 downloads2y agoHugging Face26Nunatic /dream-layout-svg-dataset-v2image1K<n<10K0 likes9 downloads2y agoHugging Face27jkanishkha0305 /text-based-layout-generation-datasetimage1K<n<10K1 likes8 downloads3y agoHugging Face28Adam789 /layout_datasetimagen<1K0 likes7 downloads3y agoHugging Face29vichetkao /table_layout_dataset_v1gated Table Dataset - Image & LabelMe & OBB Annotation (Train/Val Split) Dataset Overview Comprehensive table detection dataset with ground truth LabelMe polygon annotations and OBB (Oriented Bounding Box) data, split into training and validation sets. Total examples: 10,000 image-annotation pairs Train: 8,000 (80.0%) Validation: 2,000 (20.0%) Total size: 670.25 MB Language: km Document types: Table/Chart documents Ground truth: LabelMe polygon annotations… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/table_layout_dataset_v1.imageobject-detection10K<n<100K0 likes5 downloads4mo agoHugging Face30Nunatic /dream-layout-svg-dataset-v4Fix double boxes on symmetry places! Orz image10K<n<100K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.