iciclab/Robot_physics_interaction_with_Human_linguistic
Dual-Camera Speech and Interaction Dataset This directory contains the processed output of 50 synchronized recording sessions. Each session combines first-person video, third-person video, human speech, 30 fps image sequences, and a text label derived from the read-aloud script. For the research context behind the dataset, see PROJECT_README.md. Directory structure ACTION_NUMBER/ ├── first_person_view.mp4 ├── third_person_view.mp4 ├── audio.wav ├── text.txt ├──… See the full description on the dataset page: https://huggingface.co/datasets/iciclab/Robot_physics_interaction_with_Human_linguistic.
1584
