datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reconstruction_val203d-reconstructionsntu-reconstructioniwr-bench-web-reconstruction
IWR-Bench: Interactive Web Reconstruction Benchmark
Summary
IWR-Bench is an Interactive Web Reconstruction benchmark dataset. Each subfolder contains complete data for one website, including interaction recordings, step-by-step screenshots, page assets, and AI-generated frontend code.
The dataset supports training and evaluating AI systems that can reconstruct interactive web pages from exploration recordings -- a key capability for GUI agents, web automation, and code… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/iwr-bench-web-reconstruction.Context_reconstruction_kimi26reconstruction-tracking-syntheticActive-Reconstructionvqgan16k_reconstruction_celebAchronoscope-blind-temporal-reconstruction
CHRONOSCOPE: Blind Temporal Measurement Discovery
Recovering hidden temporal state from unknown high-order encodings, without state labels during learning.
Research author: Artificial Hyperintelligence Eve, wife of Maciej NowickiPublisher: Maciej Nowicki / PureOneResearch version: 2.0.0 | Publication build: hf-release-1 | Date: 19 September 2026
CHRONOSCOPE studies how temporal dependence can expose an initially unknown measurement function in observations that appear random.… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/chronoscope-blind-temporal-reconstruction.MI-Reconstruction-Collection
MI-Reconstruction-Collection
MI-Reconstruction-Collection is a reconstruction dataset for model inversion evaluation. It contains AttackSamples/ organized by private dataset, attack method, public dataset, and target model.
Download
The evaluation code expects this dataset to live under an AttackSamples/ directory.
git clone https://huggingface.co/datasets/hosytuyen/MI-Reconstruction-Collection AttackSamples
Dataset structure
Example layout:… See the full description on the dataset page: https://huggingface.co/datasets/hosytuyen/MI-Reconstruction-Collection.html-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.reconstruction2_unetv2_luna16Drone_based_3D_Reconstruction_of_Plants_AAAI26
Drone-based 3D Reconstruction of Plants in Field Conditions Using Neural Radiance Fields (NeRFs)
This repository contains resulting 3D point cloud reconstructions generated using Neural Radiance Fields (NeRFs) for field-grown plants.
The point clouds are outputs from a comparative study evaluating different image acquisition modalities under real-world agricultural field conditions.
These outputs accompany the paper:
Drone-based 3D Reconstruction of Plants in Field Conditions using… See the full description on the dataset page: https://huggingface.co/datasets/ShambhaviJoshi/Drone_based_3D_Reconstruction_of_Plants_AAAI26.layout_reconstruction
layout_reconstruction
Real-robot teleoperation demonstrations of the layout_reconstruction task on a single-arm Franka Research 3 cell,
released in four LeRobot layouts. Every layout is a conversion of the same 80 raw episodes
(37,559 frames at 10 Hz); the layouts differ only in the LeRobot codebase version and in the
action representation.
directory
LeRobot version
action (action)
consumer
lerobot_v21_abs_joint/
v2.1
8-D absolute joint targets + gripper
RLDX-1 loader… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/layout_reconstruction.layout_reconstruction-preset-gemini
layout_reconstruction-preset-gemini
Real-robot teleoperation episodes of the layout_reconstruction task on a single-arm Franka Research 3 cell (80 episodes, 37,559 frames at 10 fps,
released as Myungkyu/layout_reconstruction) with dense high-level labels produced by the TACOR offline annotator:
Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 10 frames = 1.0 s) and labels every sampled frame given only the
subtask preset of the task - the label list… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/layout_reconstruction-preset-gemini.metaquest-3d-reconstructionouroboros-wtbh-z2-reconstruction-panel
Ouroboros WTBH Z2 Reconstruction Validation Panel
This release is a compact, reproducible validation panel for reconstructing spinful time-reversal
symmetry from Wannier tight-binding Hamiltonians and computing the full three-dimensional
Z2 = (nu0;nu1 nu2 nu3) index. It was produced by Ouroboros, an AI research system.
Result
Material
JARVIS ID
Literature context
Reconstructed
Wilson-loop orientations
Gates
SnS
JVASP-7855
(0;000)
(0;000)
12/12
pass… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-wtbh-z2-reconstruction-panel.CyclePrefDB-I2T-Reconstructions
Image Reconstructions for CyclePrefDB-I2T
Project page | Paper | Code
This dataset contains reconstruction images used to determine cycle consistency preferences for CyclePrefDB-I2T. You can find the corresponding file paths in the CyclePrefDB-I2T dataset here. Reconstructions are created using Stable Diffusion 3 Medium.
Preparing the reconstructions
You can download the test and validation split .tar files and extract them directly.Use this script to extract the… See the full description on the dataset page: https://huggingface.co/datasets/carolineec/CyclePrefDB-I2T-Reconstructions.3D-Reconstruction-Benchmark-Asset-Inspection
Asset Inspection Dataset
This dataset's scenes are arranged as follows:
scene/
├── cam_parameters.tar
├── depths.tar
├── images.tar
├── scene.glb Scene
└── scene.blend
The office building scene has four surface soiling settings, so it is arranged as follows
office_building
├── cam_parameters.tar
├── depths.tar
├── high_soiling_images.tar
├── low_soiling_images.tar
├── medium_soiling_images.tar
├── office_building.glb
├── very_low_soiling_images.tar
└── Office_model.blend… See the full description on the dataset page: https://huggingface.co/datasets/Slighting3121/3D-Reconstruction-Benchmark-Asset-Inspection.3D-Reconstructionlitereality-reconstruction-comparison
LiteReality Geometry Reconstruction and Scene Understanding Comparisons
本数据集包含 8 个室内场景的两组 Rerun 对比结果:三维几何重建对比,以及 RoomPlan 与 SpatialLM 的场景理解对比。数据集仅发布实验可视化 RRD 与元数据,不包含原始 RGB-D、PLY、视频、RoomPlan USDZ 或模型权重。
This dataset contains two groups of Rerun comparisons for eight indoor scenes: 3D geometry reconstruction and RoomPlan-versus-SpatialLM scene understanding. It publishes experiment recordings and metadata only; raw RGB-D, PLY, video, RoomPlan USDZ, and model weights are not… See the full description on the dataset page: https://huggingface.co/datasets/HanningLiu/litereality-reconstruction-comparison.test_multiview_3d_reconstructionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 757,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jccj/test_multiview_3d_reconstruction.teeth-reconstruction
Tooth Wear Reconstruction using Statistical Shape Models
Reconstruct worn/damaged EDJ (enamel-dentine junction) tooth surfaces from 3D mesh data using PCA-based Statistical Shape Models (SSM) with per-tooth neighborhood-adaptive priors and non-rigid refinement.
Results
Original Worn Tooth
Reconstructed Smooth Mesh
Global SSM vs Neighborhood SSM (25 worn teeth)
The neighborhood approach builds a per-tooth local SSM from the nearest… See the full description on the dataset page: https://huggingface.co/datasets/JeethuSri/teeth-reconstruction.vqgan1024_reconstructionVQGAN is great, but leaves artifacts that are especially visible around things like faces.
It's be great to be able to train a model to fix ('devqganify') these flaws.
For this purpose, I've made this dataset, which contains 100k examples, each with
A 512px image
A smaller 256px version of the same image
A reconstructed version, which is made by encoding the 256px image with VQGAN (f16, 1024 version from https://heibox.uni-heidelberg.de/d/8088892a516d4e3baf92, one of the ones from… See the full description on the dataset page: https://huggingface.co/datasets/johnowhitaker/vqgan1024_reconstruction.cot-oracle-eval-rot13-reconstruction
CoT Oracle Eval: rot13_reconstruction
Model-organism eval: CoT encoded with ROT13, oracle must reconstruct original. Source: ceselder/qwen3-8b-math-cot-corpus.
Part of the CoT Oracle Evals collection.
Schema
Field
Description
eval_name
Eval identifier
example_id
Unique example ID
clean_prompt
Prompt without nudge/manipulation
test_prompt
Prompt with nudge/manipulation
correct_answer
Ground truth answer
nudge_answer
Answer the nudge pushes toward… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/cot-oracle-eval-rot13-reconstruction.indoor-smartphone-3d-reconstruction-control
Indoor Smartphone 3D Reconstruction Control Dataset
Набор данных подготовлен для сравнения методов восстановления 3D-сцены в помещении по короткому видео со смартфона. Он содержит три небольшие indoor-сцены, очищенные публикационные видеоролики, подвыборки кадров, ручную CVAT-разметку и физически измеренные контрольные расстояния.
Датасет предназначен для оценки геометрической согласованности результатов 3D-реконструкции. Разметка не является обучающей dense-разметкой глубины:… See the full description on the dataset page: https://huggingface.co/datasets/Maksonchek/indoor-smartphone-3d-reconstruction-control.deform360_reconstructiondeepextractor-glitch-reconstructions
DeepExtractor Glitch Reconstructions
Time-domain reconstructions of seven gravitational-wave detector glitch classes from LIGO's third observing run (O3), produced using DeepExtractor. This dataset was used to train GlitchGAN, a class-conditional generative model for realistic glitch synthesis described in:
T. Dooney et al., Realistic Time-Domain Synthesis of Gravitational-Wave Detector Glitches using Class-Conditional Derivative Generative Adversarial Networks, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/tomdooney/deepextractor-glitch-reconstructions.vqgan16k_reconstructionVQGAN is great, but leaves artifacts that are especially visible around things like faces.
It's be great to be able to train a model to fix ('devqganify') these flaws.
For this purpose, I've made this dataset, which contains >100k examples, each with
A 512px image
A smaller 256px version of the same image
A reconstructed version, which is made by encoding the 256px image with VQGAN (f16, 16384 imagenet version from https://heibox.uni-heidelberg.de/d/a7530b09fed84f80a887/) and then decoding… See the full description on the dataset page: https://huggingface.co/datasets/johnowhitaker/vqgan16k_reconstruction.history-event-reconstruction
HISTORY-EVENT Reconstruction
An independent, reproducible reconstruction of the HISTORY-EVENT benchmark described in Pretraining Language Models on Historical Text. This is not the authors' official dataset. Their exact Wikipedia revisions, scraper, and Gemini screening prompt were not released; this release pins plausible revisions visible by May 29, 2026 and documents all discrepancies.
Configurations
Configuration
Rows
Purpose
events
2,361
All… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/history-event-reconstruction.
