YYYYYYibo/alfworld-experimenter-gpt5mini-sft-1k
ALFWorld Experimenter GPT-5 mini SFT 1K This dataset contains 1,000 blind GPT-5 mini reasoning demonstrations for an ALFWorld expert-prefix selection task. The intended use is to give a 7B experimenter model a structured reasoning warm start before reinforcement learning, not to treat GPT-5 mini's selected depths as ground-truth labels. Task For each ALFWorld task, the experimenter receives eight failed trajectories from a frozen Qwen2.5-7B-Instruct actor and one… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/alfworld-experimenter-gpt5mini-sft-1k.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face