qa-evaluation
mistral-updated-model-for-qa_scoresystem-evaluationsetfit_baai_wix_qa_gpt-4o_cot-few_shot_remove_final_evaluation_e1_1726759073.27929setfit_baai_wix_qa_gpt-4o_cot-instructions_remove_final_evaluation_e1_larger_train_17setfit_baai_wix_qa_gpt-4o_improved-cot_chat_few_shot_remove_final_evaluation_e1_one_osetfit_baai_wix_qa_gpt-4o_cot-instructions_remove_final_evaluation_e2_larger_train_17QA-LoRA-T5-Fine-Tune-And-Evaluation
rag-qa-evaluation-dataset
RAG QA Evaluation Dataset
Overview
This dataset contains test cases for evaluating Retrieval-Augmented Generation
(RAG) and Large Language Model (LLM) applications.
The dataset is designed from a software testing and quality engineering
perspective.
Dataset Structure
Each test case contains:
Field
Description
question
User question sent to the AI application
context
Context available to the AI application
expected_answer
Expected… See the full description on the dataset page: https://huggingface.co/datasets/gaffarshaikh07/rag-qa-evaluation-dataset.fatwa-qa-evaluation
Fatwa QA Evaluation Dataset
Dataset Description
This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers.
Dataset Statistics
Total Samples: 2,000
Average Question Length: 243.9 characters
Average Answer Length: 492.3 characters
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.pedants_qa_evaluation_bench
pedants_qa_evaluation
This dataset evaluates candidate answers for various question-answering (QA) tasks across multiple datasets such as Jeopardy!, hotpotQA, nq-open, narrativeQA, and BIOMRC, etc. See details in paper. It contains questions, reference answers (ground truth), model-generated candidate answers, and human judgments indicating whether the candidate answers are correct.
Dataset Details
Column
Type
Description
question
string
The question asked… See the full description on the dataset page: https://huggingface.co/datasets/zli12321/pedants_qa_evaluation_bench.MediLite-QA-Response-EvaluationMediLite-QA-Final-Evaluationragas_QA_evaluation_datasetThe dataset comprises question answer pairs generated by the Mistral-7B-Instruct-v0.3 model, over a sample of the documents available in the CiGi knowledge base. The generative Q-A creation for evaluating CiGi was preferred over manual annotation because it yields diverse queries whose distribution better reflects downstream user intents. The dataset presentes the question-answer pairs, along with a reference to the source snippet that was used by the model to construct the answer.
