zzzhr97/Pi-Bench
Pi-Bench Tasks Code and full benchmark: GitHub repositoryPaper: arXiv:2605.14678Project page: simplified-reasoning.github.io/Pi-Bench This lightweight dataset exposes only the task.yaml files from Pi-Bench so people can quickly inspect the benchmark tasks in the Hugging Face Dataset Viewer. Pi-Bench evaluates proactive personal assistant agents in long-horizon workflows. It contains 100 multi-turn tasks across 5 domain-specific personas: researcher, marketer, pharmacist… See the full description on the dataset page: https://huggingface.co/datasets/zzzhr97/Pi-Bench.
152
Update dataset card with repository and citation
Update README.md
Upload Pi-Bench task YAML dataset
initial commit
