datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SkillHarm
SkillHarm
Lifecycle-Aware Skill-Based Attacks via Automated Construction
📄 Paper · 🌐 Project Page · 💻 GitHub · 🤗 Data
Agent skills occupy a privileged position in the agent workflow — agents are expected to implicitly follow and execute them — which makes third-party skills a vulnerable supply-chain attack surface. SkillHarm is a benchmark of skill-based attacks across the skill-use lifecycle, paired with a systematic taxonomy of 12 skill-relevant risks. Every attack is… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/SkillHarm.SKILLRET
SkillRet Benchmark
📄 Technical report: SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents (arXiv:2605.05726)
Dataset Overview
SkillRet is a retrieval benchmark for matching natural-language user requests to agent skills. It contains a curated library of public agent skills from GitHub with synthetic training and evaluation queries.
Dataset Statistics
Metric
Value
Total Records
218,157
Total File Size
714 MB
Total… See the full description on the dataset page: https://huggingface.co/datasets/ThakiCloud/SKILLRET.Skill-3D
Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning
This repository contains the dataset splits for Skill-3D, a framework for agentic 3D spatial reasoning presented in the paper Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning.
Project Page | GitHub Repository
Dataset Description
The skill3d_splits released here are used for skill evolution, post-training (SFT/GRPO), and evaluation across four major 3D spatial reasoning… See the full description on the dataset page: https://huggingface.co/datasets/lhy-zju/Skill-3D.SkillRouter-Eval-Core
SkillRouter Eval Core
This dataset hosts the public SkillRouter evaluation benchmark outside the GitHub repository so users can download it without relying on Git LFS bandwidth.
Contents
tasks.jsonl: 87 benchmark task descriptions to be routed.
relevance.json: ground-truth skill ids, task types, and graded relevance labels.
easy/: 78,361 candidate skills in gzip-sharded JSONL.
hard/: 79,141 candidate skills in gzip-sharded JSONL.
manifest.json: file layout… See the full description on the dataset page: https://huggingface.co/datasets/pipizhao/SkillRouter-Eval-Core.SkillTrustBench
SkillTrustBench
SkillTrustBench is a benchmark dataset for evaluating security analysis of agent skills: reusable capability packages that extend an AI agent through natural-language instructions, tool-use guidance, and optional executable or reference assets. Each case follows an agent-skill-style layout, with a SKILL.md entrypoint that defines when and how the skill should be used, plus optional scripts, references, assets, configuration files, or agent definitions.
The… See the full description on the dataset page: https://huggingface.co/datasets/cuhk-zhuque/SkillTrustBench.evovling_skills
Evolving Skills Benchmark
This dataset is presented in the paper EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?.
Project page: https://mas-orchestra.salesforceresearch.ai/evoharness/
Companion to evovling_tools: instead of testing adaptation to a growing
tool universe, this benchmark tests an agent's ability to discover,
author, and reuse its own SKILL library as it works through a stream of
tasks. The same general builder method is applied to two source… See the full description on the dataset page: https://huggingface.co/datasets/ZixuanKe/evovling_skills.opendata-bodypose
SkillCorner Open Data — Body Pose
3D body-pose data derived from broadcast video, released alongside the
SkillCorner Open Data repository as a
joint initiative between SkillCorner and
PySport.
Initial testing release. Two matches, published so the community can work
with the format and tell us what is useful before we consider a wider release.
Feedback is genuinely wanted — open an issue on the
opendata repo or reply in the
Community tab here.
What is in here… See the full description on the dataset page: https://huggingface.co/datasets/SkillCorner/opendata-bodypose.minimax-m3-deepsearchqa-skill-eval
MiniMax M3 DeepSearchQA Skill Eval
Evaluates minimax/minimax-m3 on google/deepsearchqa using a Pi agent, You.com MCP tools, and a research skill optimized for this harness, model, and tool surface.
MiniMax M3 Medium Reasoning with the You.com research skill reached 74.85% adjusted F1 on DeepSearchQA, above the paper's GPT-5 High Reasoning F1 result. Public artifacts are available for inspection and reproduction.
Links
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/youdotcom/minimax-m3-deepsearchqa-skill-eval.MF-Skills🚨 Please request access with your institutional email to get access to the dataset.
MF-Skills Dataset
Project page | Paper | Code
Dataset Description
MF-Skills is a large-scale dataset for advancing expert-level music understanding and reasoning in (large) audio-language models. It builds upon audio samples from LAION-DISCO and augments them with rich metadata extracted using a suite of open-source large audio-language models (LALMs) and specialized music analysis… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/MF-Skills.SWE-Skills-BenchDataset Summary
SWE-Skills-Bench is a benchmark dataset for evaluating whether injected skill documents — structured packages of procedural knowledge — measurably improve LLM agent performance on real-world software engineering tasks.
The dataset contains 49 skills spanning 565 task instances across six software engineering domains (Deployment & DevOps, Analytics & Monitoring, API Development, Data Science & ML, Security & Testing, and Developer Tools). Each skill is grounded in an authentic… See the full description on the dataset page: https://huggingface.co/datasets/GeniusHTX/SWE-Skills-Bench.skillretrieval-data
SkillRetrieval: Curated Skill Store & Index
Pre-built skill store and FAISS index for skill-retrieval-mcp.
Download with skill-mcp pull --include-index.
Contents
File
Size
Description
processed/skills.db
6.4 MB
SQLite database, 374 skills, FTS5 enabled
indices/sentence-transformers/all-MiniLM-L6-v2/index.faiss
561 KB
FAISS IndexFlatIP over L2-normalized 384-dim vectors
indices/sentence-transformers/all-MiniLM-L6-v2/skill_ids.json
7.4 KB
Row-order… See the full description on the dataset page: https://huggingface.co/datasets/zcheng256/skillretrieval-data.skillspanThis is the SkillSpan dataset created by:
@inproceedings{zhang-etal-2022-skillspan,
title = "{S}kill{S}pan: Hard and Soft Skill Extraction from {E}nglish Job Postings",
author = "Zhang, Mike and
Jensen, Kristian and
Sonniks, Sif and
Plank, Barbara",
booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
month = jul,
year = "2022",
address =… See the full description on the dataset page: https://huggingface.co/datasets/jjzha/skillspan.av-skillsAV-Skills
Audio-visual instruction and temporally grounded reasoning data for Nemotron-Labs-Audio-Visual Flamingo
AV-Skills supports joint understanding of video, speech, sound, music, and long-range temporal context in real-world videos.
Dataset Summary
AV-Skills is the audio-visual instruction and reasoning dataset for
Nemotron-Labs-Audio-Visual Flamingo, an open audio-visual language model for long and
complex… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/av-skills.Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malay-Dialect-Instructions.skillspanThis is the SkillSpan dataset created by:
@inproceedings{zhang-etal-2022-skillspan,
title = "{S}kill{S}pan: Hard and Soft Skill Extraction from {E}nglish Job Postings",
author = "Zhang, Mike and
Jensen, Kristian and
Sonniks, Sif and
Plank, Barbara",
booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
month = jul,
year = "2022",
address =… See the full description on the dataset page: https://huggingface.co/datasets/ankita627/skillspan.skills-in-the-wild
Skills in the Wild — Open Audit of AI Agent Extensions
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/skills-in-the-wild")
The first open, reproducible audit of real agent extensions (Skills, MCP, rules files) on GitHub.
Schema
file
rows
columns
manifest.jsonl
3,168
repo, path, sha, surface, html_url
findings.jsonl
742
rule_id, severity, category, evidence
files.jsonl
3,168
n_findings, worst_severity, rule_ids… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/skills-in-the-wild.qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.SkillLifeBench
SkillLifeBench — Dataset
This directory contains the complete dataset for SkillLifeBench: Benchmarking Lifecycle Security of LLM Agent Skills (NeurIPS 2026 Datasets & Benchmarks Track).
Directory Structure
SkillLifeBench/
├── README.md # this file
├── LICENSE # CC BY 4.0
├── registry.jsonl # 194 entries (flat JSONL, for dataset viewer)
├── schema/
│ └── vuln_schema.json # JSON Schema for vulnerability… See the full description on the dataset page: https://huggingface.co/datasets/SkillLifeBench2026/SkillLifeBench.all-skills-from-skills-sh
Dataset Overview
This dataset is collected from skills.sh, a website that aggregates various command-line skills. Each skill typically includes a brief description, detailed documentation, and an installation command.
We have crawled all publicly available skill pages, resulting in approximately 40,000 to 50,000 skill records. The data is stored in a structured format, making it suitable for analysis, retrieval, or further development.
Field Descriptions
Each skill… See the full description on the dataset page: https://huggingface.co/datasets/tickleliu/all-skills-from-skills-sh.agent-skill-vulnerabilities
Agent Skill Vulnerability Scenarios (defanged, teaching)
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/agent-skill-vulnerabilities")
Deliberately-vulnerable, defanged agent-extension artifacts — for training detectors & hands-on learning.
Schema
column
meaning
id
scenario
artifact_type
SKILL.md / mcp.json
content, walkthrough
artifact + defense
Related AltaySec resources
🕵️ uncloak scanner:… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/agent-skill-vulnerabilities.atr-skill-benchmark
ATR Skill-Security Benchmark
A labeled corpus of SKILL.md files for evaluating detection of malicious agent
skills — prompt injection, tool poisoning, credential theft, malware droppers
and supply-chain attacks hidden inside natural-language agent instructions.
Published as part of Agent Threat Rules (ATR),
an open, vendor-neutral detection standard for AI agents (like Sigma, but for
agent attacks).
Why this exists
SKILL.md files are natural-language instructions… See the full description on the dataset page: https://huggingface.co/datasets/Agent-Threat-Rule/atr-skill-benchmark.skillreason-bench
SkillReason-Bench
Official code and evaluation toolkit: github.com/donghong1/SkillReason
Released inference adapter, benchmark evaluation scripts, and reproducibility
instructions are available in the repository.
SkillReason-Bench evaluates agent skill retrieval for implicit user requests.
Each request describes a concrete task goal while leaving the required skill
or execution procedure unstated. A retriever must identify the most relevant
skill from a large, heterogeneous… See the full description on the dataset page: https://huggingface.co/datasets/donghongjiang/skillreason-bench.SkillCorpus
SkillCorpus
SkillCorpus is the skill-package corpus released with SPT: Skills as Pre-Training Data for Agentic Language Models. The records contain reusable tool semantics, workflows, and supporting package metadata for agentic language-model mid-training research.
Dataset splits
The 35,411 records are shuffled with Python's random.Random(42) and allocated in a 7:2:1 ratio using the largest-remainder method.
Split
Records
Train
24,788
Validation
7,082… See the full description on the dataset page: https://huggingface.co/datasets/CSeemy/SkillCorpus.Skill2-Bench
Skill²-Bench
Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning".
Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/Skill2-Bench.agent-skill-malware
Agent Skill Malware: Malicious vs Benign Agent Skills
Binary classification dataset of OpenClaw agent skill files (SKILL.md).
127 malicious + 223 benign = 350 samples, deduplicated by content hash.
Source
Malicious: Real skills from malicious campaigns targeting ClawHub (Feb 2026), extracted from the openclaw/skills GitHub archive. They use social engineering in markdown instructions to trick agents/users into running malware -- primarily AMOS (Atomic macOS Stealer)… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/agent-skill-malware.governed-skill-evolution
Governed Skill Evolution from Persistent Agent Experience
Prospective ablation and cross-model transfer study of three experience-retention conditions for governed Agent Skill evolution: no persistent history, flat chronological history, and a persistent Pattern Registry with a forward-chained Skill Impact Ledger.
Author: Song Luo
Version: 1.0.0
Source snapshot: d717c32396cfff1bef2800296541a70e9b4cabb8
Canonical repository: rrrrrredy/governed-skill-evolution
Zenodo:… See the full description on the dataset page: https://huggingface.co/datasets/RedinGhost/governed-skill-evolution.skillops-paper
SkillOps
Paper and evaluation artifacts for SkillOps: A Practical Framework for Designing, Testing, and Operating Modular Skills in Personal AI Agents by Song Luo.
Read the full paper: Read online · PDF · Zenodo · GitHub
GitHub source: rrrrrredy/skillops-paper
Source commit: 097a6bab83a3f326333c6e1f477cfeaef7664dea
Versioned research record: Zenodo DOI 10.5281/zenodo.20907648
Concept DOI: 10.5281/zenodo.20061198
Author: Song Luo
This Hub repository packages the paper… See the full description on the dataset page: https://huggingface.co/datasets/RedinGhost/skillops-paper.pm-skills-instruct
PM Skills instruct — professional judgment as training data
Instruction-tuning data distilled from pm-claude-skills — 426 open-source SKILL.md files that codify how senior professionals produce PRDs, postmortems, exec updates, launch plans, churn analyses, and 25 professions' worth of real work artifacts. Regenerated deterministically from the repo's canonical sources on every release; the version tag here matches the repo's release tag.
What's inside
File… See the full description on the dataset page: https://huggingface.co/datasets/mohitagw/pm-skills-instruct.lemonseed-mixed-skills
lemonseed-mixed-skills
LemonSeed — mixed Go + Sudoku + addition + prose stream.
Contents
mixed.jsonl (11428 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/claude-opus-4.6-4.7-reasoning-8.7k.
