CoolFace
Datasetpublic

keplerccc/Robo2VLM-1

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Abstract Vision-Language Models (VLMs) acquire real-world knowledge and general reasoning ability through Internet-scale image-text corpora. They can augment robotic systems with scene understanding and task planning, and assist visuomotor policies that are trained on robot trajectory data. We explore the reverse paradigm - using rich, real, multi-modal robot trajectory… See the full description on the dataset page: https://huggingface.co/datasets/keplerccc/Robo2VLM-1.

sourceHugging Faceapache-2.0updated 11mo agoView on Hugging Face
20likes3.1kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
keplerccc/Robo2VLM-1 · CoolFace