interactive-benchmark
Interactive_Benchmarks
Interactive Benchmarks (IB)
Interactive Benchmarks for Evaluating Interactive Reasoning and Agent Capabilities
Usage
from datasets import load_dataset
dataset = load_dataset("interactivebench/Interactive_Benchmarks")
Citation
If you use the Interactive_Benchmarks dataset in your research, please consider citing it as follows:
@misc{interactivebench,
title={Interactive Benchmarks for Evaluating Interactive Reasoning and Agent… See the full description on the dataset page: https://huggingface.co/datasets/interactivebench/Interactive_Benchmarks.Interactive_Benchmarks
