datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pseudo-camera-10k-structured-json
pseudo-camera-10k, structured JSON captions
The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled.
The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.augmented-skin-images-50k-structuredstructured_imagesSTARE-structured-analysis-of-the-retinaarXiv:2501.18921https://arxiv.org/abs/2501.18921
structured-vitalsmetamdp-robosuite-franka-moving_ball-l3-structured-train-state16-h50-v1metamdp-robosuite-franka-moving_ball-l2-structured-train-state16-h50-v1wikipedia_structured_contentsSmolVLM_Essay_Structuredstructured_images_easyStructurEditBench_v3_filteredstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetcell_w2_structuredstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetstructured3d-archWe provide parametric architectures for use with SSTK extracted from Structure3D.
|-- arch # JSON files specifying the [architecture](https://github.com/smartscenes/sstk/wiki/Architecture-Format)
|-- arch_renders # Renderings of the architecture
|-- roomId # Renderings colored by roomId
|-- images # png images
|-- camera_poses # Information about the camera paramters used for the rendering
|-- metadata… See the full description on the dataset page: https://huggingface.co/datasets/3dlg-hcvc/structured3d-arch.structured-generation-information-extraction-vlms-openbmb-RLAIF-V-DatasetStructurEditBench_v3StructurEditBench_v3_filteredstructure_data
structure_data — unified-schema 3D reconstruction datasets
Three datasets for image-to-3D structure prediction (SFT data for Qwen3-VL-4B-Instruct):
v11canon — single Infinigen objects (chairs, tables, cabinets, etc.)
scene_v3 — multi-object Infinigen scenes
physx_v11canon — PartNet-Mobility articulated objects, rotated to v11 chirality
Layout
v11canon/
train.jsonl # 7045 rows
test_seen.jsonl # 393 rows (held-out objects in seen categories)… See the full description on the dataset page: https://huggingface.co/datasets/wenjingbian/structure_data.cottonweed-structured-captionsStructurEditBench_v4_qwenStructurEditBench_v3_filtered_labelsstructured-generation-mydatagd-frames-128x96-structuredStructurEditBench_v2_1004
