datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SkillOpt_Lite_Benchmarks
SkillOpt_Lite Benchmarks
Train / val / test splits used by the SkillOpt_Lite project.
One multi-config repo containing all six benchmarks:
Config
Rows (train / val / test)
Content shipped
searchqa
400 / 200 / 1400
Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA.
docvqa
107 / 53 / 374
Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.vistr-4dagent-skillopt-pilot
ViSTR 4d-Agent SkillOpt pilot trajectories (qwen3-vl-plus)
Teacher rollout trajectories from the pi-4d-agent SkillOpt pipeline, produced on
gaozhe's AMD MI308X machine (fork commit 18d38aa95, branch gaozhe — see
packages/4d-agent/docs/handover_for_tqh.md in the fork for environment notes).
Provenance
Teacher: qwen3-vl-plus via amap-gateway (streaming patched for cumulative
args — tool-call arguments in these trajectories are clean)
Dataset: ViSTR-Bench-Public @… See the full description on the dataset page: https://huggingface.co/datasets/MihailSlutsky/vistr-4dagent-skillopt-pilot.SkillOpt-Lite
