SkillEvolBench-Team/Skill-Evol-Bench
SkillEvolBench Dataset SkillEvolBench is a diagnostic benchmark for testing whether LLM agents can convert episodic task experience into reusable procedural skills. It accompanies the paper SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills. This Hugging Face dataset page hosts the benchmark assets used by the paper's skill-evolution protocol: role-instantiated task directories, verification assets, and curated seed skills. The Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SkillEvolBench-Team/Skill-Evol-Bench.
Remove stray __pycache__/.pyc files
Restore case-sensitive-column-trap solution; sync benchmark tasks/skills
Update GitHub repository link
Update dataset card
Add dataset card
Upload benchmark skills and tasks
initial commit
