CoolFace
Datasetpublic

OS-Copilot/OS-Shepherd-100K

OSReward Data OSReward is a multimodal reward-model dataset for judging whether GUI-agent trajectories complete the user's task. The release contains: Supervised fine-tuning data in a LLaMAFactory-compatible ShareGPT layout. Deduplicated screenshots packaged in independently extractable tar shards. Training and validation data for GRPO. The expected answer contains a brief evidence-based analysis followed by a final line in exactly one of these forms: Judge: SUCCESS Judge:… See the full description on the dataset page: https://huggingface.co/datasets/OS-Copilot/OS-Shepherd-100K.

sourceHugging Faceupdated 28d agoView on Hugging Face
4likes468downloads
13 commits on main
77d61ac28d ago

Add files using upload-large-folder tool

QiushiSun
457b8ee28d ago

Mark WebTrail supplement uploaded and refresh checksums

QiushiSun
631f72428d ago

Document WebTrail 100K increment and refresh integrity manifests

QiushiSun
bcec93928d ago

Add 6,386 WebTrail last5 SFT trajectories

QiushiSun
290330328d ago

Add files using upload-large-folder tool

QiushiSun
77295022mo ago

Normalize agent model names

QiushiSun
6159cf72mo ago

Update 2026-08-07

QiushiSun
aec66142mo ago

Remove training code and scripts

QiushiSun
43486182mo ago

Add a complete SFT dataset preview

QiushiSun
bc61e882mo ago

Configure SFT preview and remove dataset counts from README

QiushiSun
43e03fb2mo ago

Add files using upload-large-folder tool

QiushiSun
a9cf38e2mo ago

Add files using upload-large-folder tool

QiushiSun
a9ecae42mo ago

initial commit

QiushiSun