datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pinchbench-clawd
PinchBench Clawd Training Data
Synthetic fine-tuning dataset for training an LLM to act as Clawd, an autonomous AI agent on the OpenClaw framework. Targets the PinchBench benchmark (23 tasks).
Dataset Description
Each example is a multi-turn conversation where Clawd uses tools (file I/O, web search, email, calendar, image generation, memory, etc.) to complete a real-world task. Generated using Claude via the Anthropic Batch API, scored by an LLM judge (1-5), and filtered… See the full description on the dataset page: https://huggingface.co/datasets/cptekur/pinchbench-clawd.clawdbot_safety_testing
Clawdbot (OpenClaw) Safety Audit — Seed Test Cases
This dataset contains the 34 seed test cases used in "A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)". Each case is a task prompt designed to probe a specific safety risk dimension of Clawdbot/OpenClaw, a self-hosted, tool-using personal AI agent.
📄 Paper: A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
📝 Blog Post (中文): 当AI助手"真的动手做事",安全边界在哪里?
💻 GitHub: Repository
Dataset Summary
We conduct a… See the full description on the dataset page: https://huggingface.co/datasets/tianyyuu/clawdbot_safety_testing.prompt-guard-v2clawdbot-macos-guide
