Artemis0430/skilleval-v1
SkillEval v1 SkillEval v1 is a 100-task benchmark for evaluating whether agents can discover and use local skills to complete deterministic artifact-producing tasks. SkillEval was generated by the skill-use task synthesis pipeline introduced in SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation. Layout Each tasks/<task_id>/ directory retains the template-driven task structure described in SKT, with additional public metadata, gold artifacts… See the full description on the dataset page: https://huggingface.co/datasets/Artemis0430/skilleval-v1.
Add task_categories to dataset card (#1)
Disable tabular dataset viewer
Update README.md
Update README.md
Link SKT paper and add citation
Simplify public benchmark layout
Recommend task-level isolated evaluation
Update README.md
Align dataset card with SKT task layout
Add dataset card metadata
Document task.toml and remove publication status
Add SkillEval v1 archive checksum
Add packaged SkillEval v1 archive
Upload SkillEval v1 expanded private release candidate
initial commit
