datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
skillsbenchWarning: The leaderboard above is generated by Hugging Face eval-results and may be incomplete until evaluation_framework: benchflow is accepted and deployed. The audited SkillsBench v1.1 result archive is https://huggingface.co/datasets/benchflow/skillsbench-leaderboard, with all retained submissions normalized under submissions/skillsbench/v1.1/ and compact official exports under leaderboard/skillsbench/v1.1/.
Warning: The dataset is a read-only mirror. The primary source for this benchmark… See the full description on the dataset page: https://huggingface.co/datasets/benchflow/skillsbench.SkillArena-datasets
SkillArena Offline Datasets
Offline evaluation data for SkillArena — a validated automatic benchmark generation framework for AI agent skills, targeting NeurIPS 2026 Datasets & Benchmarks Track.
Overview
This dataset provides domain-specific task input data for 289 AI agent skills across 13 domains. Each skill has 50 curated data files designed as meaningful agent task inputs — files an agent could receive and act upon (analyze, transform, validate, or generate from). The… See the full description on the dataset page: https://huggingface.co/datasets/JiaaqiLiu/SkillArena-datasets.deepseek-v2-codder-minecraft-apiSkillFlow-Task
Dataset Card for SkillFlow Test Tasks
Dataset Summary
SkillFlow Test Tasks is the task repository used in the SkillFlow benchmark for evaluating lifelong skill discovery, skill revision, and cross-task procedural transfer in autonomous agents.
The dataset contains 166 runnable tasks organized into 20 workflow families spanning five broad domains:
Finance & Economics
Operations & Supply Chain
Healthcare & Life Sciences
Governance & Strategy
Data & Document Intelligence… See the full description on the dataset page: https://huggingface.co/datasets/zhang-ziao/SkillFlow-Task.skilleval-v1
SkillEval v1
SkillEval v1 is a 100-task benchmark for evaluating whether agents can discover
and use local skills to complete deterministic artifact-producing tasks.
SkillEval was generated by the skill-use task synthesis pipeline introduced in
SKT: Skill-Use Training at Scale via Verified Synthetic Data
Generation.
Layout
Each tasks/<task_id>/ directory retains the template-driven task structure
described in SKT, with additional public metadata, gold artifacts… See the full description on the dataset page: https://huggingface.co/datasets/Artemis0430/skilleval-v1.SkillOpt_Lite_Benchmarks
SkillOpt_Lite Benchmarks
Train / val / test splits used by the SkillOpt_Lite project.
One multi-config repo containing all six benchmarks:
Config
Rows (train / val / test)
Content shipped
searchqa
400 / 200 / 1400
Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA.
docvqa
107 / 53 / 374
Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.SWE-Skills-BenchDataset Summary
SWE-Skills-Bench is a benchmark dataset for evaluating whether injected skill documents — structured packages of procedural knowledge — measurably improve LLM agent performance on real-world software engineering tasks.
The dataset contains 49 skills spanning 565 task instances across six software engineering domains (Deployment & DevOps, Analytics & Monitoring, API Development, Data Science & ML, Security & Testing, and Developer Tools). Each skill is grounded in an authentic… See the full description on the dataset page: https://huggingface.co/datasets/GeniusHTX/SWE-Skills-Bench.agent-skills
Index — Lucio's Agent Skills & Projects
Each project now lives in its own repo, so you get its full README, its own licence, and its own version history. This page is just the map.
This repo also keeps a full snapshot of every skill for anyone who wants them all in one download — see the Files and versions tab. The individual repos below are the canonical ones.
Agent Skills
Skill
What it does
Licence
relic
Portable AI personality & memory, in pure… See the full description on the dataset page: https://huggingface.co/datasets/LucioLiu/agent-skills.claude-agent-skills-benchmark
Claude Agent Skills Benchmark
Claude Agent Skills 评测数据集
Description
A benchmark dataset for evaluating whether LLMs can accurately trigger and execute domain-specific Skills on the Claude Code platform. Skills are designed by vertical domain experts with varying complexity levels (based on attachments: scripts, references, assets, and reference markdown files).
Evaluation Scenarios Cover:
Office automation, coding, investment promotion, financial services, industrial… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/claude-agent-skills-benchmark.Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malay-Dialect-Instructions.qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.skill-diffs
skill-diffs
Commit-by-commit revision history of agent skills (SKILL.md files) scraped from public GitHub repos. Each record is a (before, after, intent) tuple capturing how a skill was iteratively refined through human feedback.
v0.5 covers 4 platforms — Anthropic Claude, OpenClaw, OpenCode, and Hermes Agent — with PR title/body metadata as richer intent labels, MinHash + semantic clustering for dedup, structural diff_summary for filtering by edit type, aggregate quality_score for… See the full description on the dataset page: https://huggingface.co/datasets/shl0ms/skill-diffs.astra-skills
🧰 ASTRA Skills: Agent Skill Tool-use Repository Atlas
A large-scale atlas of real-world AI agent skills discovered from public GitHub repositories
ASTRA Skills (Agent Skill Tool-use Repository Atlas) is a deduplicated corpus of agent skill directories centered on SKILL.md, collected with the Astra crawler for research on tool use, agent instruction design, and skill retrieval.
Quick Start ·
At a Glance ·
Directory Format ·… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/astra-skills.agent-skills
Agent Skills Dataset
61,650 agent skills collected from GitHub repositories. Each skill contains name, description, and full markdown content. Useful for skill retrieval and agent training.
Dataset Structure
id: Unique identifier for the skill
name: Skill name
description: Skill description and usage instructions
owner: Repository owner
repo: Repository name
skill_md_path: Path to the skill markdown file
content: Full skill content in markdown format
Usage… See the full description on the dataset page: https://huggingface.co/datasets/LittleDinoC/agent-skills.multifaceted-skill-of-mind
Dataset Card for Multifaceted Skill-of-Mind 🧠🤹
🤖 Thanos-1B | 🤖 Thanos-3B | 🤖 Thanos-8B | 💻 Github | 📄 Arxiv | 📕 PDF
🚨 Disclaimer: All models and dataset are intended to be used for research purposes only.
Dataset Summary
Multifaceted Skill-of-Mind Dataset is the first publicly available skill-of-mind-annotated dialogue dataset, encompassing multi-turn, multifaceted conversational skills along with explanations across various interactive scenarios (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/passing2961/multifaceted-skill-of-mind.scentience-olfaction-skills
scentience-skills
Agent skill definitions for olfaction-aware AI — part of the Scentience olfaction platform.
Provides enhanced connectivity to Scentience olfactory processing units (OPUs).
Built to the agentskills.io open standard. Compatible with Claude Code, OpenAI Codex, and Google Antigravity.
Skills
Skill
Purpose
ble-device
Connect to OPU hardware via BLE; sample or stream sensor data
olfactory-navigation
Plume tracking and chemical source… See the full description on the dataset page: https://huggingface.co/datasets/kordelfrance/scentience-olfaction-skills.Skill2-Bench
Skill²-Bench
Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning".
Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/Skill2-Bench.SkillGenBench
SkillGenBench
SkillGenBench is a benchmark for evaluating LLM skill generation from explicit repository- and document-grounded corpora. Each benchmark instance exposes visible generation materials and an instance-specific evaluation bundle. The official v1 task set contains 187 enabled tasks across three source types: 123 Code Repo tasks, 28 Code Doc tasks, and 36 Domain Knowledge Doc tasks.
Repository Layout
data/task_manifest.csv: the Hugging Face-loadable task index.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-skillgenbench/SkillGenBench.open-skills
Open Skills: Complete skills.sh Archive
133,149 agent skills from 8,808 publishers on skills.sh
What is this?
A full dump of skills.sh as a single Parquet file. Every skill listed on the site has been collected into this dataset: README content, install commands, weekly install counts, GitHub stars, security audit results, and per-platform install breakdowns. If it's on skills.sh, it's in here.
The file is sorted by weekly installs (most popular first) and compressed… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-skills.robot-dog-skills
ShadowPEFT Dataset
This repository contains the dataset used in the paper ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning.
ShadowPEFT is a centralized parameter-efficient fine-tuning (PEFT) framework that performs layer-level refinement through a depth-shared shadow module. This dataset includes dialogue samples used for tasks such as robot intent generation to evaluate the effectiveness of the ShadowPEFT framework across different configurations.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/shadow-llm/robot-dog-skills.skillchainbench
SkillChainBench Dataset Archive
This archive contains the submitted SkillChainBench benchmark data for the NeurIPS 2026 E&D track. The executable code is distributed separately in SkillChainBench_Code.zip.
Contents
benchmark/episodes/factorized_final_v3/: 60 submitted synthetic benchmark episodes.
benchmark/skills/: 10 submitted skill manifests.
workdir_seeds/skillchain_seed_clean_noepisodes_v3/: clean agent-visible workspace seed used by the main non-oracle protocol.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-skillchainbench/skillchainbench.Reasoning-Skill
TRS Reasoning Skill Cards
News (2026-04-18):This paper has been accepted to ACL 2026 (Oral).
Project links: arXiv · GitHub · Hugging Face Dataset · Interactive Demo
This folder is a HuggingFace-ready staging package for the skill-card data used by Thinking with Reasoning Skills and its camera-ready additions.
The package includes compact skill-card corpora where they are already small/sanitized, plus manifest-only entries for larger local sources. It intentionally… See the full description on the dataset page: https://huggingface.co/datasets/stallone0000/Reasoning-Skill.pm-skills-instruct
PM Skills instruct — professional judgment as training data
Instruction-tuning data distilled from pm-claude-skills — 426 open-source SKILL.md files that codify how senior professionals produce PRDs, postmortems, exec updates, launch plans, churn analyses, and 25 professions' worth of real work artifacts. Regenerated deterministically from the repo's canonical sources on every release; the version tag here matches the repo's release tag.
What's inside
File… See the full description on the dataset page: https://huggingface.co/datasets/mohitagw/pm-skills-instruct.skills-2m
🧠 Skills-2M: GitHub-scale Agent Skills Info Atlas
🔎 A GitHub-scale metadata index of 2M+ agent skills for researchers studying the agent-skill ecosystem
Skills-2M is a large-scale SQLite index of 2M+ agent skill records collected from GitHub, normalized to help researchers study skill discovery, repository structure, metadata patterns, retrieval, and corpus construction.
Quick Start ·
At a Glance ·
Schema ·
Queries ·
Responsible Use… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/skills-2m.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/claude-opus-4.6-4.7-reasoning-8.7k.SkillFlow-Dataset
SkillFlow Dataset
This repository stores the IID training and validation data used by the SkillFlow training code.
Code
The training code is available at:
https://github.com/beita6969/SkillFlow
Files
File
Split
Samples
train_v3.json
train
3500
test_iid_v3.json
iid validation
798
Paper alignment
This release is aligned with the in-distribution benchmark families described in the SkillFlow appendix: HotpotQA, TriviaQA… See the full description on the dataset page: https://huggingface.co/datasets/beita6969/SkillFlow-Dataset.Skill2-Bench
Skill²-Bench
Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning".
Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/yinghuihe/Skill2-Bench.blended-skill-talk
Blended Skill Talk
Dataset Summary
This dataset contains conversations between two personas with additional context previous utterances free messages guided messages suggestions and guided chosen suggestions allowing for the creation of natural multi-modal conversations with personality empathy and knowledge
The conversations are designed to measure a full range of technical competencies such as dialogue flow management including response times topic control and… See the full description on the dataset page: https://huggingface.co/datasets/anezatra/blended-skill-talk.skill-based-medium-terminus2-sft
skill-based-medium-terminus2-sft
Full SFT dataset for terminal-agent fine-tuning, derived from
nvidia/Nemotron-Terminal-Corpus
(config skill_based_medium), converted to the terminus-2 "thinking-preservation"
chat format and reproducibly shuffled. 89,343 multi-turn agent trajectories across
11 terminal skills, ready to train with the AReaL
SFT recipe in ethanewer/posttraining-2606.
This is the dataset used by config_terminus2_l40s_default.yaml in that repo. Pair it
with the base… See the full description on the dataset page: https://huggingface.co/datasets/eewer/skill-based-medium-terminus2-sft.Skillex-research
Skillex Research
Skillex Research is a deterministic, training-ready transformation of
Anthropic/enabling-independent-research
for the Gram/Skillex backend. It converts the upstream privacy-preserving aggregate cluster tables into:
clusters: a normalized evidence and retrieval view with stable identifiers, explicit attribution,
confidence intervals, sparse cross-facet signals, and deterministic train/validation/test splits;
gram_skillex: paired supervised examples for… See the full description on the dataset page: https://huggingface.co/datasets/Cenedril/Skillex-research.
