datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgentChat-Test
Test Set Description
This directory contains the test set used for tool-use evaluation. The JSON files under Test-JSON/ are organized by task type:
SingleTaskProcessing/tool-select_test.json: single-tool selection tasks.
ParallelProcessing/parallel-call_test.json: parallel tool-call tasks.
ProactiveSeeking/searchTools_test_predictions_kept.json: proactive tool-search tasks.
TaskDecomposition/muti-tool-select_test.json: multi-tool task decomposition tasks.… See the full description on the dataset page: https://huggingface.co/datasets/leungtianle/AgentChat-Test.agentic-asr
Agentic ASR
Public consolidated audio and ASR result dataset for the OSWorld and
WildClawBench benchmark families.
Layout
osworld/: synthetic raw/colloquial speech, human recordings, DNS-noise
pairs, task images, ASR results, and reports.
wildclawbench/: 60 formal colloquialized prompts, synthetic speech,
20 synthetic ASR condition tables, and ten-participant human recordings.
task0_template derivatives are excluded.
metadata/conditions.jsonl: model, variant… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/agentic-asr.amazing_audio
