Rosie-Lab/compas3d
CoMPAS3D: A Dataset and Benchmark for Interactive Motion CoMPAS3D (Complex Multi-Level Person-Interaction Annotated Salsa Dataset) is a large-scale motion capture dataset designed to support research on nonverbal, physical communication through dance. It contains over 3 hours of improvised salsa duet performances by 18 dancers across beginner, intermediate, and professional skill levels. Each sequence features high-fidelity 3D motion data in the form of SMPL-X (.npz) files… See the full description on the dataset page: https://huggingface.co/datasets/Rosie-Lab/compas3d.
CoMPAS3D: A Dataset and Benchmark for Interactive Motion
<!-- Provide a quick summary of the dataset. -->
CoMPAS3D (Complex Multi-Level Person-Interaction Annotated Salsa Dataset) is a large-scale motion capture dataset designed to support research on nonverbal, physical communication through dance. It contains over 3 hours of improvised salsa duet performances by 18 dancers across beginner, intermediate, and professional skill levels. Each sequence features high-fidelity 3D motion data in the form of SMPL-X (.npz) files, along with video visualizations and synchronized audio (.mp4). We provide detailed frame-level annotations of move types, stylistic variation, and execution errors for 50% of the sequences (.txt).
Collected in a controlled studio using a 20-camera Vicon system at 120fps, the dataset includes 72 long-form sequences (each 2.5 minutes) of naturalistic leader-follower interaction. Participants were drawn from diverse dance backgrounds, and all provided informed consent for data release. CoMPAS3D offers a rich testbed for studying embodied dialogue, social fluency, and multi-agent coordination in AI and robotics.
Dataset Details
This dataset contains motion capture of improvised salsa dance by 18 individuals forming 9 pairs. The dances were improvised to 4 songs, with 2 takes recorded per song.
Dataset Description
<!-- Provide a longer summary of what this dataset is. -->
- Curated by: Rosie Lab, Simon Fraser University
- Language(s): English
- License: CC-BY-NC-4.0, except for audio in .mp4 files where rights retained by owners
Dataset Sources [optional]
<!-- Provide the basic links for the dataset. -->
- Paper: Bermet Burkanova, Yasaman Etesam, Payam Jome Yazdian, Trinity Evans, Chuxuan Zhang, Zoe Stanley, Paige Tuttösí, Angelica Lim. "CoMPAS3D: A Dataset and Benchmark for Interactive Motion." https://arxiv.org/abs/2507.19684
Uses
<!-- Address questions around how the dataset is intended to be used. -->
This dataset was created to encourage research in socially interactive embodied AI and creative, expressive humanoid motion generation. As salsa contains a move vocabulary and implict grammar rules, we propose tasks that mirror those in spoken language processing:
- Solo Motion Segmentation, Classification, Transcription (Automatic Speech Recognition)
- Solo Motion Generation (Speech Synthesis)
- Follower Motion Generation (Listener Speech Synthesis)
- Pair Motion Generation and Analysis (Conversation Synthesis)
- Style Transfer (Proficiency Adaptation)
Direct Use
<!-- This section describes suitable use cases for the dataset. --> A long term goal is to develop salsa dancing humanoids that can safely and creatively dance with each other and with real humans, adapting to their partner's proficiency, using haptic signaling as a primary form of communication. This dataset was developed for research use.
Out-of-Scope Use
This dataset is not to be used for commercial purposes.
<!--## Dataset Structure
This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. -->
Dataset Creation
Curation Rationale
<!-- Motivation for the creation of this dataset. --> Salsa is “arguably the world’s most popular partnered social dance form" and offers a challenging testbed for humanoid embodied interaction algorithms. Communication between the leader and follower is almost entirely haptic, signaled by subtle pushes and pulls.
Data Collection and Processing
<!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. -->
The data was recorded using a Vicon Motion Capture System, comprising 20 Vero MoCap cameras recording at 120 FPS within a capture volume of approximately 72 cubic meters. Participants wore Vicon motion capture suits equipped with 53 reflective markers using Vicon's "FrontWaist" configuration. The resulting .c3d files were converted using MOSH into SMPL-X (.npz files). Using witness camera recordings, the resulting SMPL-X renderings were synchronized with the audio tracks (.mp4 files)
Who are the source data producers?
<!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. -->
Participants were recruited from a university salsa dance club as well as a professional Latin dance school. Skill levels were defined by dance experience: beginners (3-6 months), intermediates (1-3 years), and professionals (over 4 years).
Annotation process
<!-- This section describes the annotation process such as annotation tools used in the process, the amount of data annotated, annotation guidelines provided to the annotators, interannotator statistics, annotation validation, etc. -->
The annotations were completed using ELAN. Segmentation involved marking the start and end frames of each 8-count dance sequence, typically corresponding to a complete dance move, based solely on the musical rhythm. Annotation was done using 4 annotation tracks: paired move labels, individual dancer move and styling annotations, and error classification.
Who are the annotators?
<!-- This section describes the people or systems who created the annotations. --> The annotator was a salsa expert with 15 years of salsa dance experience and competitive judging experience.
Personal and Sensitive Information
<!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. -->
The dataset does not contain personally identifiable data. However, the SMPL-X meshes provide estimates of body shape and sex.
Bias, Risks, and Limitations
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
This dataset contains LA-style salsa dance, one of the most popular variants of salsa. Other variants exist, such as New York style and Cuban style salsa. Contact information was generated using post-processing of mesh intersections after 1-cm mesh inflations and do not reflect touch recorded from raw sensors.
Recommendations
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
When using this dataset for salsa dance generation or move recognition, users should consider the variant of salsa, similar to considering varied dialects or accents in English or other natural languages.
Citation
<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. --> Bermet Burkanova, Yasaman Etesam, Payam Jome Yazdian, Trinity Evans, Chuxuan Zhang, Zoe Stanley, Paige Tuttösí, Angelica Lim. "CoMPAS3D: A Dataset and Benchmark for Interactive Motion." https://arxiv.org/abs/2507.19684
BibTeX:
@misc{burkanova2026compas3ddatasetbenchmarkinteractive,
title={CoMPAS3D: A Dataset and Benchmark for Interactive Motion},
author={Bermet Burkanova and Yasaman Etesam and Payam Jome Yazdian and Trinity Evans and Chuxuan Zhang and Zoe Stanley and Paige Tuttösí and Angelica Lim},
year={2026},
eprint={2507.19684},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2507.19684},
}Glossary of Moves, Styling, and Errors
<!-- If relevant, include terms and calculations in this section that can help readers understand the dataset or dataset card. -->
Acknowledgments
We would like to thank Giorgio Becherini and Dr. Michael Black for their assistance in MOSH conversion to SMPL-X format. We also thank Ahmet Tasel and Jim Su for their help in learning the motion capture process and initial discussions. This work would also not be possible without support from the Rajan Family.
Dataset Card Details and Contact
This dataset was collected as part of the M.Sc. thesis of Bermet Burkanova under the supervision of Angelica Lim.
