ruotong-pan/CAGB
Credibility-aware Generation Benchmark (CAGB) is a benchmark constructed to evaluate the credibility-aware generation ability of models, dealing with flawed information in the context. This benchmark encompasses the following three specific scenarios where the integration of credibility is essential: Open-domain QA 2WikiMultiHopQA HotpotQA Musique RGB Time-sensitive QA EvoTempQA Misinformation Polluted QA NewsPollutedQA Read our paper for more insights on credibility-aware generation.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face