datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.AspectSim-Evaluation-Benchmark
Dataset Card for AspectSim
Dataset Details
Dataset Description
AspectSim is a large-scale aspect-conditioned document-pair similarity evaluation benchmark. Each instance consists of two full documents, a natural-language aspect on which the comparison is based, and a human-interpretable similarity label on an ordinal scale. The benchmark spans five diverse domains: news, opinion, hotel reviews, medical literature, and scientific peer reviews, enabling… See the full description on the dataset page: https://huggingface.co/datasets/aspectsim/AspectSim-Evaluation-Benchmark.
