datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cad-gen-freecad-bench
Parametric CAD Bench — results dataset
Run-by-run results for Parametric CAD Bench, a benchmark that
measures whether AI agents can author editable FreeCAD models from
natural-language part descriptions. 1000 rows, one per
(agent, model, task_id, trial) over the
gnucleus-ai/cad-bench@v1
task suite. The public leaderboard view of this data lives at
cadbench.ai.
What's in here
data/cad-bench-v1.parquet — the row table. Each row carries the
composite + sub-scores… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench.cad-gen-freecad-bench-v2
Parametric CAD Bench v2 — results dataset
Run-by-run results for Parametric CAD Bench v2, a benchmark that measures
whether AI agents can author editable FreeCAD models from natural-language part
descriptions. This archive contains 1,000 rows: one trial for each of 100 tasks
across the 10 public jobs on the live
gnucleus-ai/cad-bench@v2
leaderboard.
What's in here
data/cad-bench-v2.parquet — the trial index. Each row carries the
continuous reward and its… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench-v2.
