leopoldmaillard/sceneteract-grpo
SceneTeract GRPO Training Set Action-level feasibility samples for post-training a VLM against a geometric verifier. Each row is one atomic interaction — an image, a prompt, and a label that was measured rather than annotated — ready to drop into TRL's GRPOTrainer. 8,073 samples over 1,132 3D-FRONT living rooms and dining rooms and three agent profiles. from datasets import load_dataset ds = load_dataset("leopoldmaillard/sceneteract-grpo") ds["train"] # 6,473 samples / 905… See the full description on the dataset page: https://huggingface.co/datasets/leopoldmaillard/sceneteract-grpo.
096
