datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
java_evaluation_benchmarksbenchmark-evaluation-resultsfactual-consistency-evaluation-benchmarkThis is a mix of 22 datasets that were used to evaluate factual consistency models available in this collection.
The distribution is available below:
subset
count
halueval_cnndm
19998
alisawuffles/WANLI
5000
Seahorse
4135
ExpertQA
3702
fib_xsum
3534
anli
3200
scitail
2126
Lfqa
1911
DeFacto
1836
llm_summaries_cnndm
1829
llm_summaries_xsum
1726
Reveal
1705
FactCheck-GPT
1565
FoolMeTwice
1379
aggrefact_xsum
1335
ClaimVerify
1087
aggrefact_cnndm
1017… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/factual-consistency-evaluation-benchmark.
