ruotong-pan/CAGB
Credibility-aware Generation Benchmark (CAGB) is a benchmark constructed to evaluate the credibility-aware generation ability of models, dealing with flawed information in the context. This benchmark encompasses the following three specific scenarios where the integration of credibility is essential: Open-domain QA 2WikiMultiHopQA HotpotQA Musique RGB Time-sensitive QA EvoTempQA Misinformation Polluted QA NewsPollutedQA Read our paper for more insights on credibility-aware generation.
Credibility-aware Generation Benchmark (CAGB) is a benchmark constructed to evaluate the credibility-aware generation ability of models, dealing with flawed information in the context. This benchmark encompasses the following three specific scenarios where the integration of credibility is essential:
- Open-domain QA
- 2WikiMultiHopQA
- HotpotQA
- Musique
- RGB
- Time-sensitive QA
- EvoTempQA
- Misinformation Polluted QA
- NewsPollutedQA
Read our paper for more insights on credibility-aware generation.
