datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SSR-3DFRONT
SSR-3DFRONT: Structured Scene Representation for 3D Indoor Scenes
This dataset provides a processed version of the 3D-FRONT dataset with structured scene representations for text-driven 3D indoor scene synthesis and editing.
Mor information about ReSpace: http://respace.mnbucher.com
For detailed usage instructions, training details, and examples, see the associated repository: https://github.com/GradientSpaces/respace
Our model weights for SG-LLM:… See the full description on the dataset page: https://huggingface.co/datasets/gradient-spaces/SSR-3DFRONT.3D-SynthPlace_indoor_scenes_dataset
3D-SynthPlace indoor scene dataset in OptiScene (NeurIPS2025)
This is the 3D-SynthPlace dataset in the paper OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization (NeurIPS2025).
3D-SynthPlace dataset JSON File Format Specification
The basic format of the scene description is JSON. This format is used to describe room floor and interior object layouts. You can refer prompts_all_scenes.json to… See the full description on the dataset page: https://huggingface.co/datasets/B3rrYang/3D-SynthPlace_indoor_scenes_dataset.OpenSCAD_3D_SFT
OpenSCAD 3D-SFT Model Card
This model card documents the dataset schema, prompt design, distributional composition, and training configuration underlying the OpenSCAD Supervised Fine-Tuning (SFT) model. The model is designed to synthesize valid, compilation-ready, and parametric OpenSCAD source code from natural-language specifications provided in either Chinese or English.
Dataset Overview
The corpus comprises synthetically generated SFT dialogues, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/Chunjiang-Intelligence/OpenSCAD_3D_SFT.toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001
ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 001
This dataset repo records the exact local training-data state visible to the dynamic epoch launcher.
It intentionally stores manifests and audit records rather than duplicating large Parquet shards.
Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_001_special_structure_current_step_002000.pt
Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT
Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001.toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002
ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 002
This dataset repo records the exact local training-data state visible to the dynamic epoch launcher.
It intentionally stores manifests and audit records rather than duplicating large Parquet shards.
Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_002_special_structure_delta_step_002750.pt
Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT
Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002.fineweb-edu10B-superbpe50304-3digit
FineWeb-Edu 10BT, tokenized with a two-phase 50,304-token SuperBPE vocabulary (3-digit number splitting)
All of HuggingFaceFW/fineweb-edu
sample-10BT, pre-tokenized in the binary shard format used by
modded-nanogpt, with a custom SuperBPE
vocabulary trained with BatchBPE (v2) on the
same corpus.
Phase 1 of the vocabulary split numbers into groups of up to 3 digits (\p{N}{1,3}). This is
the baseline in a planned comparison of first-phase digit-splitting rules, with sibling… See the full description on the dataset page: https://huggingface.co/datasets/alexandermorgan/fineweb-edu10B-superbpe50304-3digit.3d-prompt3-digit-arithmetic-scratchpad-traces
Contents
Split
Rows
train
100,000
validation
4,000
test
4,000
total
108,000
Splits are prompt-disjoint — no expression appears in more than one split,
and commutative swaps and trace keys are de-duplicated across splits to prevent
split leakage.
Operation
Rows
×
32,000
÷
32,000
+
22,000
−
22,000
Operands lie in [−999, 999]. Division answers use a fixed DDD.ddd form
(round-half-up to three decimals); division by zero is an atomic <nan>.… See the full description on the dataset page: https://huggingface.co/datasets/vmal/3-digit-arithmetic-scratchpad-traces.3dnews-articles
Dataset Card for 3DNews Articles
Dataset Summary
The dataset comprises news articles from the Russian technology website 3DNews, covering the period from 2003 to 2024. It covers the latest updates in the world of digital technology and insightful commentary from industry experts, spanning the years 2003 to 2024.
Languages
The dataset is mostly in Russian, but there may be other languages present.
Dataset Structure
Data Fields
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/3dnews-articles.3D_Synthetic_Petroleum_Derived_GEMS
3D Synthetic Petroleum-Derived GEMS (MVP Release)
📌 Dataset Overview
This dataset contains 197 elite, highly complex 3D molecular structures derived from petroleum fractions. Designed specifically for petrochemicals, materials science, organic semiconductors, and specialized additives, these compounds represent a curated "Golden Fund" of stable, complex hydrocarbons.
Unlike drug-like molecules, this dataset focuses on polycyclic architectures, rigid 3-ring… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/3D_Synthetic_Petroleum_Derived_GEMS.my-distiset-3d6680f8
Dataset Card for my-distiset-3d6680f8
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Bruno2023/my-distiset-3d6680f8/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Bruno2023/my-distiset-3d6680f8.synthetic-preference-data
Synthetic Preference Data
A small, synthetically generated preference dataset intended for testing
RLHF / DPO training pipelines. Each example contains a prompt and two
responses — one correct (chosen), one subtly flawed (rejected).
Generation Procedure
Generator model: gpt-4o-mini (OpenAI)
Generation method: OpenAI Structured Outputs (response_format=PreferenceExample) — guarantees schema-valid JSON.
Prompt templates: 4 (factual, step-by-step reasoning, technical how-to… See the full description on the dataset page: https://huggingface.co/datasets/antony-bryan-3D2Y/synthetic-preference-data.3dprintAdvanced Business Model:
Company Name: QuickPrint
Objective: Establish a profitable and scalable 3D printing service catering to local businesses, individuals, and educational institutions while leveraging emerging technologies.
Initial Investment: $25,000 (seed funding)
Projected Revenue Streams:
Service Fees: Offer various services such as:
Prototyping
Production parts for industries like aerospace or automotive.
Product Sales: Sell printed products like:
Phone cases and laptop sleeves with… See the full description on the dataset page: https://huggingface.co/datasets/Mi6paulino/3dprint.my-distiset-3d1aa117
Dataset Card for my-distiset-3d1aa117
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/trendfollower/my-distiset-3d1aa117/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/trendfollower/my-distiset-3d1aa117.
