datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
burpwn-usage
burpwn Usage — fine-tuning dataset (CLI + MCP tool-use)
An instruction-tuning dataset that teaches an LLM to operate
burpwn — a transparent intercepting
proxy and rootless sandbox for AI-driven web pentesting on Linux — across three
interfaces: instruction-style CLI prose, real Bash tool calls (how an
agent actually runs burpwn from a CLI session / under the PreToolUse hook), and
the MCP (Model Context Protocol) tool interface. Roughly half the records are
genuine multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/own2pwn-fr/burpwn-usage.SO-Python_QA-API_Usage-tanh_score
Stack Overflow Python Q&A Dataset
Description
Filtered Python Q&A with API_Usage subcategory without:
Images
Links
Blocks of code
Scores in Q1-Q3 scaled with MaxAbsScaler. Tanh function applyed to joint Scores.
agent-tool-usage
Agent Tool Usage Dataset
Goal–tool–input–output samples for training tool-using AI agents.
