claude-skill
claude-agent-skills-benchmark
Claude Agent Skills Benchmark
Claude Agent Skills 评测数据集
Description
A benchmark dataset for evaluating whether LLMs can accurately trigger and execute domain-specific Skills on the Claude Code platform. Skills are designed by vertical domain experts with varying complexity levels (based on attachments: scripts, references, assets, and reference markdown files).
Evaluation Scenarios Cover:
Office automation, coding, investment promotion, financial services, industrial… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/claude-agent-skills-benchmark.claude-skills
Claude Skills Dataset
This dataset contains curated SKILL.md files plus generated structured summaries and embeddings.
Columns
name: skill name (from SKILL.md frontmatter)
description: short description (from SKILL.md frontmatter)
full_content: full SKILL.md content (includes frontmatter metadata)
repo: source repository
split: dataset split label
qwen3emb_description: embedding vector for description (float list)
gpt_domain: concise domain label extracted from the skill… See the full description on the dataset page: https://huggingface.co/datasets/huzey/claude-skills.claude-skills-diff
claude-skills-diff
This dataset contains Git history snapshots for SKILL.md files from repos in huzey/claude-skills.
What is a row?
Each row corresponds to a (repo, skill_path) pair where the content changed a lot between its first and last commit for that path.
We export:
text_initial / initial_sha / initial_date
text_final / final_sha / final_date
optional text_middle / middle_sha / middle_date
diffs: diff_initial_final and if middle exists: diff_initial_middle… See the full description on the dataset page: https://huggingface.co/datasets/huzey/claude-skills-diff.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/claude-opus-4.6-4.7-reasoning-8.7k.european-countries
European Countries
A small reference dataset of European countries with capital, population,
area, currency, ISO codes, and EU membership status.
Population and area figures are approximate recent estimates. Russia and
Turkey are transcontinental; full territory figures are reported.
mcp-conversation
mcp-conversation
This dataset was created using the Claude Dataset Skill.
