CoolFace
Datasetpublic

ruotong-pan/CAGB

Credibility-aware Generation Benchmark (CAGB) is a benchmark constructed to evaluate the credibility-aware generation ability of models, dealing with flawed information in the context. This benchmark encompasses the following three specific scenarios where the integration of credibility is essential: Open-domain QA 2WikiMultiHopQA HotpotQA Musique RGB Time-sensitive QA EvoTempQA Misinformation Polluted QA NewsPollutedQA Read our paper for more insights on credibility-aware generation.

sourceHugging Facemitupdated 2y agoView on Hugging Face
1likes166downloads
Dataset Card

Credibility-aware Generation Benchmark (CAGB) is a benchmark constructed to evaluate the credibility-aware generation ability of models, dealing with flawed information in the context. This benchmark encompasses the following three specific scenarios where the integration of credibility is essential:

  • —Open-domain QA
  • —2WikiMultiHopQA
  • —HotpotQA
  • —Musique
  • —RGB
  • —Time-sensitive QA
  • —EvoTempQA
  • —Misinformation Polluted QA
  • —NewsPollutedQA

Read our paper for more insights on credibility-aware generation.