datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SQuADDS_Layouts
SQuADDS Layouts - versioned GDS artifacts for superconducting quantum hardware
SQuADDS Layouts is the geometry-artifact companion to
SQuADDS_DB, the
Superconducting Qubit And Device Design and Simulation Database. It provides
checksum-verified GDS files, stable geometry identities, and machine-readable
geometry metadata so a simulation result can be traced to the exact layout
that produced it.
Homepage: https://lfl-lab.github.io/SQuADDS/
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/SQuADDS/SQuADDS_Layouts.SQuADDS_Layout_Embeddings
SQuADDS Layout Embeddings
Versioned layout representations for the 24,106 GDS artifacts in
SQuADDS/SQuADDS_Layouts.
Static embedding model v0
static-embedding-v0 implements the original SQuADDS proof-of-concept model:
v0 = parameter_sum + geometric_moments + flattened_shape_bitmap
Each unit-normalized vector has 9,227 dimensions:
Block
Dimensions
Contents
Parameter sum
1
Permutation- and parameter-count-invariant sum of numerical design options… See the full description on the dataset page: https://huggingface.co/datasets/SQuADDS/SQuADDS_Layout_Embeddings.persian-ocr-community-dataset-layout
Persian OCR Community Layout Annotations
Resumable layout annotations for the page images in
Reza2kn/persian-ocr-community-dataset.
Each row points to an exact source dataset revision, Parquet shard, blob, and row. It includes the
page identifier, page dimensions, handwriting flag, and structured layout boxes produced by
datalab-to/surya_layout2 at confidence threshold
0.4.
The boxes field contains label, confidence, raster-order position, and pixel coordinates
x0, y0, x1, y1.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-community-dataset-layout.DocVQA_layoutLM
Dataset Card for "DocVQA_layoutLM"
More Information needed
adresyze-ad-layouts
AdResyze Ad Layout Annotations (v2)
Layout annotations for 302 Indian brand advertisements: which region of an ad is the
logo, the headline, the call to action, the product, the price, the background — and
where each one sits.
Built for AdResyze, which uses these
layouts to reflow a single creative into every platform aspect ratio without squashing
the logo or cropping the call to action.
Try it: builditwithgk--adresyze-ui.modal.run
· Code:… See the full description on the dataset page: https://huggingface.co/datasets/builditwithgk/adresyze-ad-layouts.DocVQA_for_LayoutLM
Dataset Card for "DocVQA_layoutLM_large"
More Information needed
CIVQA-TesseractOCR-LayoutLM
CIVQA TesseractOCR LayoutLM Dataset
The Czech Invoice Visual Question Answering dataset was created with Tesseract OCR and encoded for the LayoutLM.
The pre-encoded dataset can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR
All invoices used in this dataset were obtained from public sources. Over these invoices, we were focusing on 15 different entities, which are crucial for processing the invoices.
Invoice number
Variable… See the full description on the dataset page: https://huggingface.co/datasets/SpringRollMonster/CIVQA-TesseractOCR-LayoutLM.alwas-analog-layout-dataset
ALWAS Analog Layout Dataset
Synthetic dataset for training ML models in the ALWAS (Analog Layout Workflow Automation System) pipeline.
Dataset Description
4,000 analog IC layout blocks with complete metadata, stage transitions, and labels for:
Hours estimation — actual vs estimated hours
Complexity classification — Low / Medium / High
Bottleneck risk prediction — Low / Medium / High
Completion time prediction — stage-by-stage transition history
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/muthuk1/alwas-analog-layout-dataset.rollout_smolvla_so101_lv2_single_cube_to_box_random_layout__green_sync_0723_1027_20260723_102833This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/geonmin-kim/rollout_smolvla_so101_lv2_single_cube_to_box_random_layout__green_sync_0723_1027_20260723_102833.so101-fixed-layout-3camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101",
"total_episodes": 53,
"total_frames": 13250,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:53"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/igor-saprygin/so101-fixed-layout-3cam.xarm6-pick-mustard-bottle-sim-v4-layoutsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "xarm6",
"total_episodes": 50,
"total_frames": 7694,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/crislmfroes/xarm6-pick-mustard-bottle-sim-v4-layouts.SO101-lv2-single-cube-to-box-random-layout-grasp-correction-v1so101-fixed-layout-vlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101",
"total_episodes": 50,
"total_frames": 12500,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/igor-saprygin/so101-fixed-layout-vla.so101-fixed-layout-lift-vlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101",
"total_episodes": 3,
"total_frames": 750,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/igor-saprygin/so101-fixed-layout-lift-vla.layoutlm_sqad
Dataset Card for "layoutlm_sqad"
More Information needed
blender-3d-layout-720playoutlmv3-document-qa-v2alayoutlmv3-document-qaCIVQA_EasyOCR_LayoutLM_Validation
CIVQA EasyOCR LayoutLM Validation Dataset
The CIVQA (Czech Invoice Visual Question Answering) dataset was created with EasyOCR, and it is encoded for LayoutLM models. This dataset contains only the validation split. The train part of the dataset can be found on this URL: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_TrainThe pre-encoded validation dataset can be found on this link:… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_Validation.layout_test_202605131211019226robomme-vus-green-blue-varied-layout-20260804CIVQA_EasyOCR_LayoutLM_Train
CIVQA EasyOCR LayoutLM Train Dataset
The CIVQA (Czech Invoice Visual Question Answering) dataset was created with EasyOCR, and it is encoded for LayoutLM models. This dataset contains only the train split. The validation part of the dataset can be found on this URL: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_ValidationThe pre-encoded train dataset can be found on this link:… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_Train.layout_test_202602250712549410
