OS-Copilot/OS-Shepherd-100K
OSReward Data OSReward is a multimodal reward-model dataset for judging whether GUI-agent trajectories complete the user's task. The release contains: Supervised fine-tuning data in a LLaMAFactory-compatible ShareGPT layout. Deduplicated screenshots packaged in independently extractable tar shards. Training and validation data for GRPO. The expected answer contains a brief evidence-based analysis followed by a final line in exactly one of these forms: Judge: SUCCESS Judge:… See the full description on the dataset page: https://huggingface.co/datasets/OS-Copilot/OS-Shepherd-100K.
Add files using upload-large-folder tool
Mark WebTrail supplement uploaded and refresh checksums
Document WebTrail 100K increment and refresh integrity manifests
Add 6,386 WebTrail last5 SFT trajectories
Add files using upload-large-folder tool
Normalize agent model names
Update 2026-08-07
Remove training code and scripts
Add a complete SFT dataset preview
Configure SFT preview and remove dataset counts from README
Add files using upload-large-folder tool
Add files using upload-large-folder tool
initial commit
