zzzhr97/Pi-Bench
Pi-Bench Tasks Code and full benchmark: GitHub repositoryPaper: arXiv:2605.14678Project page: simplified-reasoning.github.io/Pi-Bench This lightweight dataset exposes only the task.yaml files from Pi-Bench so people can quickly inspect the benchmark tasks in the Hugging Face Dataset Viewer. Pi-Bench evaluates proactive personal assistant agents in long-horizon workflows. It contains 100 multi-turn tasks across 5 domain-specific personas: researcher, marketer, pharmacist… See the full description on the dataset page: https://huggingface.co/datasets/zzzhr97/Pi-Bench.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face