datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
java_evaluation_benchmarksAudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.benchmark-evaluation-resultsagent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.toxicity_benchmark-evaluationgovreport_evaluation_benchmarkfactual-consistency-evaluation-benchmarkThis is a mix of 22 datasets that were used to evaluate factual consistency models available in this collection.
The distribution is available below:
subset
count
halueval_cnndm
19998
alisawuffles/WANLI
5000
Seahorse
4135
ExpertQA
3702
fib_xsum
3534
anli
3200
scitail
2126
Lfqa
1911
DeFacto
1836
llm_summaries_cnndm
1829
llm_summaries_xsum
1726
Reveal
1705
FactCheck-GPT
1565
FoolMeTwice
1379
aggrefact_xsum
1335
ClaimVerify
1087
aggrefact_cnndm
1017… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/factual-consistency-evaluation-benchmark.AspectSim-Evaluation-Benchmark
Dataset Card for AspectSim
Dataset Details
Dataset Description
AspectSim is a large-scale aspect-conditioned document-pair similarity evaluation benchmark. Each instance consists of two full documents, a natural-language aspect on which the comparison is based, and a human-interpretable similarity label on an ordinal scale. The benchmark spans five diverse domains: news, opinion, hotel reviews, medical literature, and scientific peer reviews, enabling… See the full description on the dataset page: https://huggingface.co/datasets/aspectsim/AspectSim-Evaluation-Benchmark.Crab-role-playing-evaluation-benchmark
📄 Paper
|
📄 Github
💬 Role-playing Model
|
💬 Role-palying Evaluation Model
💬 Training Dataset
|
💬 Evaluation Benchmark
|
💬 Annotated Role-playing Evaluation Dataset
|
💬 Human-preference Dataset
1. Introduction
This is the dataset used for evalauating a role‑playing LLM.
More details can be seen at GitHub and Crab… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/Crab-role-playing-evaluation-benchmark.benchmark-evaluation-meter
Shaer-AI/benchmark-evaluation-meter
This dataset extends Shaer-AI/benchmark-evaluation by adding BiLSTM-based meter probability distributions.
Added columns
*_meter_dist_json: JSON with:
pred
pred_prob
probs (probability per meter label)
