szang18/s3-language-following-v7big
S3 Language-Following v7big Dual-arm tabletop pick corpus for the S3 language-following benchmark. 5,531 episodes over 2,520 scenes (textured rendering, 4 cameras: head / front / wrist×2, 256×256 @10fps, 98 frames/episode) Task family: "pick up the apple 〈relation〉 the 〈landmark〉" — 4 spatial relations × 6 landmark objects (18 populated cells), 3 identical apples per scene, 1 distractor landmark, 8 caption phrasings per cell Balanced ~62 demos per (cell × distractor-config);… See the full description on the dataset page: https://huggingface.co/datasets/szang18/s3-language-following-v7big.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face