datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
geometric_shapes
Geometric Shapes Dataset
This dataset contains procedurally generated images of various geometric shapes with corresponding captions. It's designed for educational purposes and testing of diffusion models.
Dataset Overview
Content: 100,000 images of geometric shapes with detailed metadata
Image size: 512x512 pixels
Format: PNG images with CSV metadata
Features: Various shapes, colors, sizes, and descriptive captions
Purpose: Educational use for training and… See the full description on the dataset page: https://huggingface.co/datasets/anokimchen/geometric_shapes.fMRI-Shape
fMRI-Shape Dataset: A Component of the fMRI-3D Dataset for MinD-3D++
This repository contains the fMRI-Shape dataset, a component of the comprehensive fMRI-3D dataset introduced and utilized in the paper MinD-3D++: Advancing fMRI-Based 3D Reconstruction with High-Quality Textured Mesh Generation and a Comprehensive Dataset. This work builds upon the initial "MinD-3D" research.
The fMRI-3D dataset consists of two components: fMRI-Shape (this dataset) and fMRI-Objaverse. Both datasets… See the full description on the dataset page: https://huggingface.co/datasets/Fudan-fMRI/fMRI-Shape.ShapeNetPartShapeNet-55Original dataset page: https://github.com/Julie-tang00/Point-BERT/
atlas-24-frozen-prefix-potential-shaping
ATLAS report 24: frozen-prefix potential shaping
1. Question and links
Read this first. This data root holds the first attempt of report 24 on the campaign's old harness (verl 0.7.1): the shaped training is complete and the unshaped training stopped at step 20 with a known problem (the subsection at the end of this section). The question was rerun on the runtime of report 25 with both trainings at 40 steps; that rerun's trajectories, exports, checkpoints and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-24-frozen-prefix-potential-shaping.shapenet_npzl-shape
Code for dataset generation: https://github.com/benikm91/drawing-dataset-generator
rl__24GPU_shaped__inferredbugs-sandboxes-verifier__exp_tas_optimal_comb__40-0basic_3D_shapesshape
Shape Geometry Dataset
Synthetic graph-based centerline representations of 3D geometric motifs (pipe-like structures).
JSON Schema
dataset.json is an array of shape records. Each record:
{
"category": "arc_90",
"nodes": [[x, y, z], ...],
"edges": [[i, j], ...],
"features": {
"curvature": [0.0, 0.1, ...],
"segment_angle": [0.0, 160.5, ...]
}
}
Field
Type
Description
category
string
Shape class label (e.g. straight, arc_90, corner)
nodes… See the full description on the dataset page: https://huggingface.co/datasets/bayang/shape.mllm-shap
MLLM-SHAP experiment datasets
Curated test splits for studying Shapley-value explanations in multimodal large language models (text and audio inputs). Each configuration is a filtered, size-controlled subset built for reproducible benchmarking—not a full copy of the upstream corpora.
Configs follow the naming pattern {task}__{source} (for example single_sentence__voice_bench).
Quick load
Pin a dataset revision for reproducibility (replace REVISION with the commit hash… See the full description on the dataset page: https://huggingface.co/datasets/Pawlo77/mllm-shap.agent-coordination-shape-outcomes
Per-Task Agent Coordination Shape Outcomes
What this dataset is
When you orchestrate LLM agents, you choose a coordination shape before running anything: solve with a
single agent, fan out and vote, decompose into parallel subtasks, chain a draft through critique,
or run an orchestrator-workers topology. Which shape is best changes from task to task, but is it predictable
per task?
This dataset is the evidence to answer that. It runs all five shapes on 159 hard… See the full description on the dataset page: https://huggingface.co/datasets/soren19/agent-coordination-shape-outcomes.shapenet_segmentationIN1k256-bfl16latents_shape_dc-ae-f32c32-sana-1.0rl__24GPU_shaped__selfinstruct-naive-sandboxes-2-verified__exp_tas_optimal_comb__40-0rl__24GPU_shaped__stackexchange-overflow-sandboxes-skywork-response__exp_tas_optimal_comb__40-0seqnorm-tis-shapedrl__24GPU_shaped__swe_rebench_patched_oracle__r2egym-nl2bash-stackdev_set_v2_rl__24GPU_shaped__selfinstruct_naive_sandboxes_2_verified__exp_tas_o57316c9adev_set_v2_rl__24GPU_shaped__inferredbugs_sandboxes_verifier__exp_tas_optimal_c937091c1terminal_bench_2_rl__24GPU_shaped__selfinstruct_naive_sandboxes_2_verified__exp93c24543swebench_verified_random_100_folders_rl__24GPU_shaped__inferredbugs_sandboxes_vf2b407c9terminal_bench_2_rl__24GPU_shaped__inferredbugs_sandboxes_verifier__exp_tas_opt2ea25390ablation-pymethods2test-shapedSHAPEFILESshape-blind-dataset
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
This dataset is part of the work "Forgotten Polygons: Multimodal Large Language Models are Shape-Blind".📖 Read the Paper💾 GitHub Repository
Overview
This dataset is designed to evaluate the shape understanding capabilities of Multimodal Large Language Models (MLLMs).
Sample Usage
This dataset is designed to be used with the evaluation code provided in the GitHub Repository. To evaluate… See the full description on the dataset page: https://huggingface.co/datasets/mgolov/shape-blind-dataset.ShapeNet
ShapeNet
Paper: https://arxiv.org/pdf/1512.03012
Files in this repo:
shapenetcore_partanno_segmentation_benchmark_v0_normal.zip - This is a drop-in replacement for https://shapenet.cs.stanford.edu/media/shapenetcore_partanno_segmentation_benchmark_v0_normal.zip (which no longer exists)
shapenetcore_partanno_segmentation_benchmark_v0_normal_npz.zip - This is a larger version (double the file size) taken from https://www.kaggle.com/datasets/mitkir/shapenet/data. It includes an… See the full description on the dataset page: https://huggingface.co/datasets/cminst/ShapeNet.spectre-shapesshape-biasshapenetpart
