datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deepseek-v2-codder-minecraft-apiSkillOpt_Lite_Benchmarks
SkillOpt_Lite Benchmarks
Train / val / test splits used by the SkillOpt_Lite project.
One multi-config repo containing all six benchmarks:
Config
Rows (train / val / test)
Content shipped
searchqa
400 / 200 / 1400
Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA.
docvqa
107 / 53 / 374
Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.SWE-Skills-BenchDataset Summary
SWE-Skills-Bench is a benchmark dataset for evaluating whether injected skill documents — structured packages of procedural knowledge — measurably improve LLM agent performance on real-world software engineering tasks.
The dataset contains 49 skills spanning 565 task instances across six software engineering domains (Deployment & DevOps, Analytics & Monitoring, API Development, Data Science & ML, Security & Testing, and Developer Tools). Each skill is grounded in an authentic… See the full description on the dataset page: https://huggingface.co/datasets/GeniusHTX/SWE-Skills-Bench.qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.skill-diffs
skill-diffs
Commit-by-commit revision history of agent skills (SKILL.md files) scraped from public GitHub repos. Each record is a (before, after, intent) tuple capturing how a skill was iteratively refined through human feedback.
v0.5 covers 4 platforms — Anthropic Claude, OpenClaw, OpenCode, and Hermes Agent — with PR title/body metadata as richer intent labels, MinHash + semantic clustering for dedup, structural diff_summary for filtering by edit type, aggregate quality_score for… See the full description on the dataset page: https://huggingface.co/datasets/shl0ms/skill-diffs.agent-skills
Agent Skills Dataset
61,650 agent skills collected from GitHub repositories. Each skill contains name, description, and full markdown content. Useful for skill retrieval and agent training.
Dataset Structure
id: Unique identifier for the skill
name: Skill name
description: Skill description and usage instructions
owner: Repository owner
repo: Repository name
skill_md_path: Path to the skill markdown file
content: Full skill content in markdown format
Usage… See the full description on the dataset page: https://huggingface.co/datasets/LittleDinoC/agent-skills.multifaceted-skill-of-mind
Dataset Card for Multifaceted Skill-of-Mind 🧠🤹
🤖 Thanos-1B | 🤖 Thanos-3B | 🤖 Thanos-8B | 💻 Github | 📄 Arxiv | 📕 PDF
🚨 Disclaimer: All models and dataset are intended to be used for research purposes only.
Dataset Summary
Multifaceted Skill-of-Mind Dataset is the first publicly available skill-of-mind-annotated dialogue dataset, encompassing multi-turn, multifaceted conversational skills along with explanations across various interactive scenarios (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/passing2961/multifaceted-skill-of-mind.Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malay-Dialect-Instructions.Skill2-Bench
Skill²-Bench
Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning".
Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/Skill2-Bench.SkillGenBench
SkillGenBench
SkillGenBench is a benchmark for evaluating LLM skill generation from explicit repository- and document-grounded corpora. Each benchmark instance exposes visible generation materials and an instance-specific evaluation bundle. The official v1 task set contains 187 enabled tasks across three source types: 123 Code Repo tasks, 28 Code Doc tasks, and 36 Domain Knowledge Doc tasks.
Repository Layout
data/task_manifest.csv: the Hugging Face-loadable task index.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-skillgenbench/SkillGenBench.pm-skills-instruct
PM Skills instruct — professional judgment as training data
Instruction-tuning data distilled from pm-claude-skills — 426 open-source SKILL.md files that codify how senior professionals produce PRDs, postmortems, exec updates, launch plans, churn analyses, and 25 professions' worth of real work artifacts. Regenerated deterministically from the repo's canonical sources on every release; the version tag here matches the repo's release tag.
What's inside
File… See the full description on the dataset page: https://huggingface.co/datasets/mohitagw/pm-skills-instruct.robot-dog-skills
ShadowPEFT Dataset
This repository contains the dataset used in the paper ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning.
ShadowPEFT is a centralized parameter-efficient fine-tuning (PEFT) framework that performs layer-level refinement through a depth-shared shadow module. This dataset includes dialogue samples used for tasks such as robot intent generation to evaluate the effectiveness of the ShadowPEFT framework across different configurations.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/shadow-llm/robot-dog-skills.open-skills
Open Skills: Complete skills.sh Archive
133,149 agent skills from 8,808 publishers on skills.sh
What is this?
A full dump of skills.sh as a single Parquet file. Every skill listed on the site has been collected into this dataset: README content, install commands, weekly install counts, GitHub stars, security audit results, and per-platform install breakdowns. If it's on skills.sh, it's in here.
The file is sorted by weekly installs (most popular first) and compressed… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-skills.skills-2m
🧠 Skills-2M: GitHub-scale Agent Skills Info Atlas
🔎 A GitHub-scale metadata index of 2M+ agent skills for researchers studying the agent-skill ecosystem
Skills-2M is a large-scale SQLite index of 2M+ agent skill records collected from GitHub, normalized to help researchers study skill discovery, repository structure, metadata patterns, retrieval, and corpus construction.
Quick Start ·
At a Glance ·
Schema ·
Queries ·
Responsible Use… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/skills-2m.SkillFlow-Dataset
SkillFlow Dataset
This repository stores the IID training and validation data used by the SkillFlow training code.
Code
The training code is available at:
https://github.com/beita6969/SkillFlow
Files
File
Split
Samples
train_v3.json
train
3500
test_iid_v3.json
iid validation
798
Paper alignment
This release is aligned with the in-distribution benchmark families described in the SkillFlow appendix: HotpotQA, TriviaQA… See the full description on the dataset page: https://huggingface.co/datasets/beita6969/SkillFlow-Dataset.Skill2-Bench
Skill²-Bench
Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning".
Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/yinghuihe/Skill2-Bench.Skillex-research
Skillex Research
Skillex Research is a deterministic, training-ready transformation of
Anthropic/enabling-independent-research
for the Gram/Skillex backend. It converts the upstream privacy-preserving aggregate cluster tables into:
clusters: a normalized evidence and retrieval view with stable identifiers, explicit attribution,
confidence intervals, sparse cross-facet signals, and deterministic train/validation/test splits;
gram_skillex: paired supervised examples for… See the full description on the dataset page: https://huggingface.co/datasets/Cenedril/Skillex-research.blended-skill-talk-fixed
Compatibility Update
This repository is a compatibility-fixed version of the original Blended Skill Talk dataset.
The original dataset can be found at:
Original Hugging Face dataset: https://huggingface.co/datasets/anezatra/blended-skill-talk
This version was created to maintain compatibility with newer versions of the Hugging Face datasets library.
Changes from the Original Dataset
The following changes were made:
Removed the unused label_candidates column.… See the full description on the dataset page: https://huggingface.co/datasets/TutorialGuide/blended-skill-talk-fixed.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/claude-opus-4.6-4.7-reasoning-8.7k.SFT_DATA-openthoughts-1k_rows-main-Qwen2.5-7B-Instruct-SkillFactoryYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-openthoughts-1k_rows-main-Qwen2.5-7B-Instruct-SkillFactory",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
BF_EVAL-cd3args-Qwen2.5-1.5B-Instruct-RLThese datasets are exactly like the Evaluation datasets except the model_responses array are budget forcing rounds.
So the first response is at a maximum total context length of 4k, the second response (2nd index in the array) is a continuation of that last response up to a total of 8,192 tokens.
Column Details
Column
Description
question
The question we want the model to answer
answer
The string answer
task
The name of the task the row belongs to
prompt
The prompt… See the full description on the dataset page: https://huggingface.co/datasets/SkillFactory/BF_EVAL-cd3args-Qwen2.5-1.5B-Instruct-RL.SFT_DATA-openthoughts-1k_rows-baseline-QwQ-AnnotatedYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-openthoughts-1k_rows-baseline-QwQ-Annotated",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"
},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
EVAL-OT-Qwen2.5-7B-Instruct-QwQ-1k_rows-RL
Column Details
Column
Description
question
The question we want the model to answer
answer
The string answer
task
The name of the task the row belongs to
prompt
The prompt we will feed into the model to solve the question
model_responses
An array of strings that the model generated to answer the prompt (usually size of 4 or 34 depending on the evaluation task)
model_responses__eval_is_correct
An array aligned with model_responses containing booleans: True when… See the full description on the dataset page: https://huggingface.co/datasets/SkillFactory/EVAL-OT-Qwen2.5-7B-Instruct-QwQ-1k_rows-RL.openclaw-hermes-repo-skills
OpenClaw and Hermes Agent Repository Skills
This dataset contains repository-skill records mined from two agent/harness repositories:
https://github.com/openclaw/openclaw
https://github.com/NousResearch/hermes-agent
It was generated by repo-skills-miner: https://github.com/peytontolbert/repository-skill-miner
The dataset is intended for retrieval, routing, classification, and analysis of reusable repository skills. It is not a raw source-code dump. Each row normalizes a mined unit… See the full description on the dataset page: https://huggingface.co/datasets/PeytonT/openclaw-hermes-repo-skills.BF_EVAL-cd3args-Qwen2.5-1.5B-Instruct-SkillFactory-RLThese datasets are exactly like the Evaluation datasets except the model_responses array are budget forcing rounds.
So the first response is at a maximum total context length of 4k, the second response (2nd index in the array) is a continuation of that last response up to a total of 8,192 tokens.
Column Details
Column
Description
question
The question we want the model to answer
answer
The string answer
task
The name of the task the row belongs to
prompt
The prompt… See the full description on the dataset page: https://huggingface.co/datasets/SkillFactory/BF_EVAL-cd3args-Qwen2.5-1.5B-Instruct-SkillFactory-RL.EVAL-cd3args-Qwen2.5-1.5B-Instruct-SkillFactory-RL
Column Details
Column
Description
question
The question we want the model to answer
answer
The string answer
task
The name of the task the row belongs to
prompt
The prompt we will feed into the model to solve the question
model_responses
An array of strings that the model generated to answer the prompt (usually size of 4 or 34 depending on the evaluation task)
model_responses__eval_is_correct
An array aligned with model_responses containing booleans: True when… See the full description on the dataset page: https://huggingface.co/datasets/SkillFactory/EVAL-cd3args-Qwen2.5-1.5B-Instruct-SkillFactory-RL.SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_reflectionsYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-cd3args-ablation-Qwen2.5-1.5B-Instruct-no_reflections",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
coding-skill-real-world-needsSynthetic Data Distillation from GPT-4o mini for Latest Programming Skills Market Needs
EVAL-cd3args-Qwen2.5-7B-Instruct-RL
Column Details
Column
Description
question
The question we want the model to answer
answer
The string answer
task
The name of the task the row belongs to
prompt
The prompt we will feed into the model to solve the question
model_responses
An array of strings that the model generated to answer the prompt (usually size of 4 or 34 depending on the evaluation task)
model_responses__eval_is_correct
An array aligned with model_responses containing booleans: True when… See the full description on the dataset page: https://huggingface.co/datasets/SkillFactory/EVAL-cd3args-Qwen2.5-7B-Instruct-RL.EVAL_MATH500-OT-Qwen2.5-7B-Instruct-SkillFactory-1k_rows-RL
Column Details
Column
Description
question
The question we want the model to answer
answer
The string answer
task
The name of the task the row belongs to
prompt
The prompt we will feed into the model to solve the question
model_responses
An array of strings that the model generated to answer the prompt (usually size of 4 or 34 depending on the evaluation task)
model_responses__eval_is_correct
An array aligned with model_responses containing booleans: True when… See the full description on the dataset page: https://huggingface.co/datasets/SkillFactory/EVAL_MATH500-OT-Qwen2.5-7B-Instruct-SkillFactory-1k_rows-RL.
