datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgentLongBench
AgentLongBench Benchmark Dataset
Standardized evaluation dataset for AgentLong tasks. This directory is the
data-only companion to the agentlong_bench codebase and follows a fixed
layout so that runners can infer knowledge/history labels directly from the
path.
Summary
The dataset contains multi-round "guess-the-entity" dialogues with either:
knowledge-intensive content (Pokemon identities), or
knowledge-free masked entities.
Each JSONL file contains samples for a… See the full description on the dataset page: https://huggingface.co/datasets/ign1s/AgentLongBench.a20k-train-embeddings
