ACERobotics/ACE-Data-0
ACE-Data-0 Human-Centric Ambient Capture as Embodied Data Engine S-Lab, Nanyang Technological University, Singapore · ACE Robotics ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI. ▶ Demo video · Full story, figures, and interactive examples on the blog What this is Learning to act in the physical… See the full description on the dataset page: https://huggingface.co/datasets/ACERobotics/ACE-Data-0.
48257k
1---2pretty_name: ACE-Data-03license: other4license_name: ace-data-0-research-license5license_link: LICENSE6viewer: false7language:8 - en9task_categories:10 - robotics11 - keypoint-detection12 - video-classification13 - audio-classification14size_categories:15 - 10K<n<100K16tags:17 - video18 - audio19 - 3d20 - timeseries21 - robotics22 - embodied-ai23 - multimodal24 - egocentric25 - multi-view26 - motion-capture27 - tactile-sensing28 - human-object-interaction29 - human-scene-interaction30 - long-horizon31 - smpl-x32 - mano33 - imitation-learning34 - vision-language-action35extra_gated_heading: Request access to ACE-Data-036extra_gated_description: >-37 ACE-Data-0 contains identifiable recordings of human participants captured38 inside real homes. Access is granted to named individuals for non-commercial39 academic research only, on acceptance of the terms below.40extra_gated_prompt: >-41 **Before requesting access, please read the following.**42 43 44 ACE-Data-0 is released **exclusively for non-commercial academic research**45 under the [ACE-Data-0 Research License46 Agreement](https://huggingface.co/datasets/ACERobotics/ACE-Data-0/blob/main/LICENSE).47 By submitting this form you confirm that you have read that agreement and that48 you agree to be bound by it.49 50 51 The dataset contains **video, audio, motion, and tactile recordings of52 identifiable human participants** who volunteered and gave informed consent53 for research release. In particular, you agree that you will **not**:54 55 56 - use the data, or anything derived from it, for any commercial purpose;57 58 - redistribute, publish, or otherwise share the data with anyone who has not59 been granted access individually through this form;60 61 - attempt to identify, contact, locate, or infer private attributes of any62 participant appearing in the recordings;63 64 - use the data to develop or evaluate biometric identification, surveillance,65 or profiling systems.66 67 68 Access is granted to **you personally and is not transferable**. Every69 colleague, student, or collaborator who needs the data must submit their own70 request.71 72 73 The details you provide below are recorded as part of your licence. Submitting74 inaccurate information, or using the dataset outside the terms above, is a75 breach of the agreement and grounds for revoking your access.76 77 78 Commercial licensing, industrial collaboration, and any use outside the scope79 above are handled separately. Please open a thread on the Community tab of this80 repository rather than submitting this form.81extra_gated_fields:82 Full name: text83 Institutional email address: text84 Institution or organization: text85 Country: country86 Position:87 type: select88 options:89 - Undergraduate student90 - Master's student91 - PhD student92 - Postdoctoral researcher93 - Faculty / Principal investigator94 - Research scientist / Research engineer95 - label: Other96 value: other97 Homepage, Google Scholar, or lab page: text98 Which parts of ACE-Data-0 do you intend to use?:99 type: select100 options:101 - Egocentric video102 - Exocentric video103 - Human body and hand motion104 - Object meshes and 6-DoF poses105 - Audio106 - Tactile107 - Annotations only108 - The complete dataset109 Describe your intended research use in at least two sentences (project, tasks, and expected outputs): text110 I confirm that I am requesting access for non-commercial academic research only: checkbox111 I have read and agree to the ACE-Data-0 Research License Agreement: checkbox112 I will not redistribute the dataset or any part of it, and I will direct colleagues to submit their own request: checkbox113 I will not attempt to identify participants, nor use the data for biometric identification, surveillance, or profiling: checkbox114 I agree to cite ACE-Data-0 in any publication or public artifact that uses it: checkbox115 I understand that access is personal, non-transferable, and may be revoked: checkbox116extra_gated_button_content: Submit access request117---118 119<div align="center">120 <h1>ACE-Data-0</h1>121 <h3>Human-Centric Ambient Capture as Embodied Data Engine</h3>122 <p>123 <b>S-Lab, Nanyang Technological University, Singapore</b>124 · 125 <b>ACE Robotics</b>126 </p>127 <p>128 <a href="https://ace-data-engine.github.io/ACE-Data-0/"><img src="https://img.shields.io/badge/Blog-ACE--Data--0-0054A6?logo=googlechrome&logoColor=white" alt="Blog"></a>129 <a href="https://arxiv.org/abs/2607.28625"><img src="https://img.shields.io/badge/arXiv-2607.28625-B31B1B?logo=arxiv&logoColor=white" alt="Technical report on arXiv"></a>130 <a href="./LICENSE"><img src="https://img.shields.io/badge/License-Research_Only-A33B20" alt="Research-only license"></a>131 <a href="https://ace-data-engine.github.io/ACE-Data-0/"><img src="https://img.shields.io/badge/Data_Files-Coming_Soon-6B7280" alt="Data files coming soon"></a>132 </p>133 <p>134 <img src="https://ace-data-engine.github.io/ACE-Data-0/assets/images/data-teaser.webp" width="100%" alt="ACE-Data-0 teaser: table-scale and room-scale ambient capture with synchronized multi-modal streams">135 </p>136 <p>137 <b>ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI.</b>138 </p>139 <p>140 ▶ <a href="https://ace-data-engine.github.io/ACE-Data-0/assets/videos/teaser-video.mp4">Demo video</a>141 · 142 <a href="https://ace-data-engine.github.io/ACE-Data-0/">Full story, figures, and interactive examples on the blog</a>143 </p>144</div>145 146## What this is147 148Learning to act in the physical world requires more than observing what an action looks like: models149must capture how first-person perception, whole-body motion, dexterous manipulation, object state,150sound, and touch evolve **together** as humans pursue goals over time. Existing datasets fragment151this experience across viewpoints, modalities, or spatial scales.152 153ACE-Data-0 records all of it in one pass, on one clock, in one world frame. Participants receive154**goal-level** instructions ("prepare a cup of tea and serve it at the table") rather than155step-by-step scripts, so planning, hesitation, and improvisation enter the data by themselves. Human156states, object states, and contact are metrically tracked or directly sensed, so the annotations are157**measured rather than estimated**: they stay correct under furniture occlusion, extreme viewpoints,158and motion blur, where image-based detectors fail.159 160| | |161| --- | ---: |162| Recorded activity | 150 hours+ |163| Video frames | 17M+ |164| Interaction episodes | 75,000+ |165| Task categories | 200+ |166| Participants | 50+ |167| Capture environments | 2 |168| Views per moment | 8+ exocentric, plus 4 egocentric fisheye |169| Take length | minutes, not seconds; up to 20-30 min for long-horizon chains |170 171These values describe the planned release and will be verified in the final release manifest.172 173## What each take contains174 175Every take shares one timeline and one world coordinate frame, and ships its own calibration and176sync tables as data. Any tracked 3D point can therefore be projected onto any pixel of any view, and177any two streams paired at any instant, without rerunning any part of the capture pipeline.178 179| Modality | Contents |180| --- | --- |181| Egocentric video | 4 fisheye views @ 20 FPS, IMU, per-frame 6-DoF headset pose from the tracked rig |182| Exocentric video | 8 synchronized views @ 30 FPS, each with intrinsics and world-frame pose |183| Human motion | 41-joint skeletons, articulated hand poses, converted SMPL-X parameters |184| Object state | scanned or 2DGS meshes, 6-DoF pose @ 60 Hz, 2D/3D boxes in all views, motion trails |185| Audio | multi-source, from the exocentric cameras and the headset |186| Tactile | full-palm pressure grids, normalized and baseline-corrected |187| Language | per-segment activity descriptions, take goal, and its sequence of sub-goals |188 189Takes come in three families: **atomic HOI** (1-3 household tasks, ~3 min), **chains of HOI** (one190continuous activity of ~20-30 min ending with the scene tidied back into order), and **human-scene191interaction** (whole-body motion and furniture contact, almost no objects, ~5 min).192 193### Room-scale directory structure194 195```text196Room-scale/197├── Body_shape/198├── calibration/199└── Take-xxxxxx/200 ├── motions/201 ├── raw/202 ├── timeline/203 ├── videos_and_annotations/204 └── caption/205```206 207- `Body_shape/` contains the body-shape beta parameters for each participant.208- `calibration/` contains calibration data for the OptiTrack system and ZED cameras, together with209 the scripts used to calculate the Homie calibrations.210- Each `Take-xxxxxx/` directory contains the data for one take:211 - `motions/` contains human and object motion data.212 - `raw/` contains the raw egocentric, exocentric, and motion-capture data.213 - `timeline/` contains the synchronized timeline for all modalities. Its `annotation.json` file214 provides `video_keep_interval_s`, which specifies the video intervals to retain after removing215 footage captured only for synchronization.216 - `videos_and_annotations/` contains the captured egocentric and exocentric videos at their217 original resolution, along with annotated videos for the human body, hands, and objects.218 - `caption/` contains a text description of the take. Caption timestamps refer to the filtered219 timeline after applying `video_keep_interval_s`.220 221### Table-scale directory structure222 223```text224Table-scale/225├── HOI/226│ ├── data/227│ │ └── chunk-000/228│ ├── meta/229│ │ └── takes/230│ ├── objs/231│ ├── raw/232│ │ ├── calibration/233│ │ ├── camera_array_aniposelib/234│ │ ├── captions/235│ │ ├── hand_poses/236│ │ ├── imu/237│ │ ├── obj_file/238│ │ ├── obj_poses/239│ │ ├── optitrack_to_camera_world_alignment/240│ │ └── poses/241│ └── videos/242│ ├── egocap/243│ └── gopro/244└── Tactile/245 └── ACE-Data-Tactile-{0..10}/246 ├── data/247 │ └── chunk-000/248 ├── meta/249 │ └── takes/250 ├── raw/251 │ ├── calibration/252 │ ├── camera_array_aniposelib/253 │ ├── captions/254 │ ├── imu/255 │ └── tactile/256 └── videos/257 ├── egocap/258 └── gopro/259```260 261- `HOI/` contains table-scale human-object interaction recordings focused on fine-grained262 hand-object manipulation.263 - `data/` contains LeRobot-compatible parquet chunks. Rows are canonical reference samples on a264 fixed-rate timeline; `frame_index` is an episode-local reference-sample index, while the raw-row265 ranges point to records in `raw/`.266 - `meta/` contains take-level metadata, including the `takes/` index and task information.267 - `objs/` contains the object mesh assets referenced by the object-tracking data.268 - `raw/` contains the captured sensor records and sidecars, including camera calibration,269 camera-array processing data, captions, human and object poses, IMU data, object files, and270 OptiTrack-to-camera alignment data.271 - `videos/` contains the source video streams, separated into `egocap/` for egocentric capture and272 `gopro/` for exocentric GoPro cameras.273 274- `Tactile/` contains the tactile table-scale release split into eleven shards,275 `ACE-Data-Tactile-0/` through `ACE-Data-Tactile-10/`. Each shard follows the same layout:276 - `data/` contains the canonical timeline samples and parquet chunks for that shard.277 - `meta/` contains take-level metadata under `takes/`.278 - `raw/` contains calibration, camera-array processing data, captions, IMU records, and raw279 tactile measurements under `tactile/`.280 - `videos/` contains the corresponding `egocap/` and `gopro/` video streams.281 282The `data/`, `raw/`, and `videos/` directories are complementary: the parquet data provides timeline283indices and raw-row ranges, sensor records remain in `raw/`, and the video bytes remain in `videos/`.284 285## How it was captured286 287| | Table-scale | Room-scale |288| --- | --- | --- |289| Space | ~30 m² desk workspace | ~200 m² furnished apartment |290| Target | fine-grained hand-object manipulation | whole-body activity and locomotion |291| Exocentric RGB | 8 × GoPro at 0.3-0.5 m | 8 × ZED One, at least 4 views on any point |292| Optical mocap | 16 × OptiTrack PrimeX 22 | 12 × OptiTrack PrimeX 22 |293| Hand pose | triangulated from 8 exo views, manually refined | Manus mocap gloves @ 60 Hz |294 295Participants wear an ACE-Ego-Head-V02 Lite headset (4 fisheye cameras, IMU, 5 tracked markers), a29641-marker mocap suit, and full-palm tactile gloves.297 298Two numbers carry the credibility of everything above. **Temporal:** all devices are registered to299the OptiTrack 60 Hz clock by photographing a nanosecond-resolution clock displayed on the mocap300host, giving millisecond-level residuals, within a single mocap frame. **Spatial:** an ArUco board301with retroreflective corners bridges exocentric cameras that share no field of view (median302reprojection error < 3 px), while the headset is solved by hand-eye calibration against its tracked303rig (~2 px), so egocentric camera poses are measured rather than estimated and do not drift.304 305The blog and the technical report cover the capture protocol, calibration, and annotation pipeline306in full.307 308## Benchmark309 310We hold out 10 hours as a test set and evaluate 30+ published methods across three levels: tactile311inference from video, human motion recovery, and hand motion from egocentric and exocentric views.312Existing methods degrade sharply under contact, occlusion, egomotion, and long horizons. In313particular, strong per-frame pose accuracy does not imply an accurate world-frame trajectory, and314egomotion, not finger articulation, dominates the error in egocentric hand reconstruction. Full315tables are in the technical report.316 317## Access318 319This repository is **gated**. Access is granted to named individuals for **non-commercial academic320research only**, and takes effect once you accept the terms on the access form above.321 322- Use your **institutional** email and describe your intended research. What you submit is recorded323 as part of your licence; inaccurate information is a breach of the agreement.324- Access is **personal and non-transferable**. Collaborators and students must each submit their own325 request. Redistributing the data terminates your licence and revokes your access.326- For commercial licensing or industrial collaboration, open a thread on the Community tab instead327 of submitting the form.328 329> **Data files are not published yet.** Repository layout, storage requirements, checksums, and330> loading examples will be added at release time. Because the streams are large, the release will331> use sharded archives rather than direct browser downloads. Approved users keep their access.332 333## Intended and prohibited uses334 335Intended for non-commercial academic research on embodied perception, human and hand motion336recovery, human-object and human-scene interaction, egocentric and multi-view video understanding,337cross-modal learning across vision, motion, audio, and touch, and imitation learning, world models,338and vision-language-action systems.339 340**Not** for identifying, re-identifying, or profiling participants; biometric recognition or341surveillance; inferring sensitive personal attributes; any commercial purpose; or representing all342homes, cultures, bodies, abilities, or household practices without further validation. The343[LICENSE](./LICENSE) is binding and defines the full set of restrictions.344 345All participants volunteered and signed informed consent covering data collection and research346release, including the appearance of their faces. **The dataset contains identifiable individuals.**347If you are a participant and want your recordings withdrawn, reach us through the Community tab and348the affected takes will be removed from subsequent releases.349 350## Limitations351 352Two sites only, so limited variation in layouts, furnishings, and lighting. Tracked objects must be353scanned and marked in advance, and state changes of articulated mechanisms, fluids, and deformable354materials are not annotated. The mocap suit, gloves, headset, and markers are visible in the355recordings and may introduce dataset-specific visual cues.356 357## License358 359[ACE-Data-0 Research License Agreement](./LICENSE): non-commercial academic research only, no360redistribution, no re-identification. Read it in full before requesting access.361 362## Citation363 364```bibtex365@article{cao2026acedata0,366 title = {ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine},367 author = {Cao, Yukang and Xie, Haozhe and Wen, Beichen and Yao, Runmao and368 Liu, Yinghao and Huang, Yue and Liao, Zhichao and Wang, Yunxiang and369 Liu, Haiheng and Tian, Xingshun and Su, Dawei and Zhuo, Long and370 Tao, Dacheng and Wang, Xiaogang and Pan, Liang and Liu, Ziwei},371 journal = {arXiv preprint arXiv:2607.28625},372 year = {2026}373}374```375 376## Contact377 378Project updates on the [blog](https://ace-data-engine.github.io/ACE-Data-0/). Questions about379access, licensing, or annotations belong on the Community tab of this repository.380 