datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatRAG-Bench
ChatRAG Bench
ChatRAG Bench is a benchmark for evaluating a model's conversational QA capability over documents or retrieved context. ChatRAG Bench are built on and derived from 10 existing datasets: Doc2Dial, QuAC, QReCC, TopioCQA, INSCIT, CoQA, HybriDialogue, DoQA, SQA, ConvFinQA. ChatRAG Bench covers a wide range of documents and question types, which require models to generate responses from long context, comprehend and reason over tables, conduct arithmetic calculations, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/ChatRAG-Bench.ChatRAG-Hi
Dataset Description:
The ChatRAG-Hi (Hindi ChatRAG Bench) dataset is based on the English version of the ChatRAG Bench, which comprises the following ten datasets: Doc2Dial, QuAC, QReCC, INSCIT, HybriDialogue, DoQA, and ConvFinQA. The dataset was translated using GCP, and approximately 500 samples were filtered from each of these sets based on backtranslation accuracy to eliminate poor translations.
The evaluation steps are described here.
Other Hindi benchmark datasets: [… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/ChatRAG-Hi.
