CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gaffarshaikh07 /rag-qa-evaluation-dataset RAG QA Evaluation Dataset Overview This dataset contains test cases for evaluating Retrieval-Augmented Generation (RAG) and Large Language Model (LLM) applications. The dataset is designed from a software testing and quality engineering perspective. Dataset Structure Each test case contains: Field Description question User question sent to the AI application context Context available to the AI application expected_answer Expected… See the full description on the dataset page: https://huggingface.co/datasets/gaffarshaikh07/rag-qa-evaluation-dataset.1 likes50 downloads5d agoHugging Face02SahmBenchmark /fatwa-qa-evaluation Fatwa QA Evaluation Dataset Dataset Description This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers. Dataset Statistics Total Samples: 2,000 Average Question Length: 243.9 characters Average Answer Length: 492.3 characters Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.tabularquestion-answering1K<n<10K0 likes44 downloads9mo agoHugging Face03zli12321 /pedants_qa_evaluation_bench pedants_qa_evaluation This dataset evaluates candidate answers for various question-answering (QA) tasks across multiple datasets such as Jeopardy!, hotpotQA, nq-open, narrativeQA, and BIOMRC, etc. See details in paper. It contains questions, reference answers (ground truth), model-generated candidate answers, and human judgments indicating whether the candidate answers are correct. Dataset Details Column Type Description question string The question asked… See the full description on the dataset page: https://huggingface.co/datasets/zli12321/pedants_qa_evaluation_bench.textquestion-answering10K<n<100K1 likes32 downloads2y agoHugging Face04Cowboygarage /MediLite-QA-Response-Evaluationtext1K<n<10K0 likes20 downloads11mo agoHugging Face05Cowboygarage /MediLite-QA-Final-Evaluationtext1K<n<10K0 likes14 downloads11mo agoHugging Face06CGIAR /ragas_QA_evaluation_datasetgatedThe dataset comprises question answer pairs generated by the Mistral-7B-Instruct-v0.3 model, over a sample of the documents available in the CiGi knowledge base. The generative Q-A creation for evaluating CiGi was preferred over manual annotation because it yields diverse queries whose distribution better reflects downstream user intents. The dataset presentes the question-answer pairs, along with a reference to the source snippet that was used by the model to construct the answer. text1K<n<10K0 likes4 downloads2mo agoHugging Face07emirMb /RAG-EVALUATION-QAtabularn<1K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.