skill
Datasets
All datasets matching “skill”skillsbench-leaderboard
SkillsBench Leaderboard and Evidence Archive
This repository stores public SkillsBench submissions, raw BenchFlow trial artifacts, trajectory evidence, audit reports, and the release-aligned official leaderboard exports.
Official benchmark definition: benchflow/skillsbenchLatest public benchmark release: SkillsBench v1.1Latest source commit: 27738384b1df694ea2ae466e416f476e94d8fab9
Current Official Release
The latest public results are under:… See the full description on the dataset page: https://huggingface.co/datasets/benchflow/skillsbench-leaderboard.skillsbench-trend-anomaly-causal-inferenec-taskskilltrainbench-public
skilltrainbench training tasks
The training half of the skilltrainbench benchmark suite: for each of the
four datasets, the dev_task_names of its pinned train/test split, in Harbor
task format.
The held-out/test tasks are not in this repository. Neither are the
published splits that are not the pin, nor the tasks that fall outside each
pinned split's task set. Use this repository for skill formation and
training; evaluate on the held-out half, which stays in the private source… See the full description on the dataset page: https://huggingface.co/datasets/armin-aptura/skilltrainbench-public.blended_skill_talk
Dataset Card for "blended_skill_talk"
Dataset Summary
A dataset of 7k conversations explicitly designed to exhibit multiple conversation modes: displaying personality, having empathy, and demonstrating knowledge.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 38.11 MB
Size of the generated dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ParlAI/blended_skill_talk.Skill-Evol-Bench
SkillEvolBench Dataset
SkillEvolBench is a diagnostic benchmark for testing whether LLM agents can convert episodic task experience into reusable procedural skills. It accompanies the paper SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills.
This Hugging Face dataset page hosts the benchmark assets used by the paper's skill-evolution protocol: role-instantiated task directories, verification assets, and curated seed skills. The Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SkillEvolBench-Team/Skill-Evol-Bench.skillsbenchWarning: The leaderboard above is generated by Hugging Face eval-results and may be incomplete until evaluation_framework: benchflow is accepted and deployed. The audited SkillsBench v1.1 result archive is https://huggingface.co/datasets/benchflow/skillsbench-leaderboard, with all retained submissions normalized under submissions/skillsbench/v1.1/ and compact official exports under leaderboard/skillsbench/v1.1/.
Warning: The dataset is a read-only mirror. The primary source for this benchmark… See the full description on the dataset page: https://huggingface.co/datasets/benchflow/skillsbench.
