datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Medical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.Competence-Based-Evaluation
Competence-Based Evaluation (Invariance Benchmark)
A benchmark for testing whether language models give the same answer to
semantically equivalent reformulations of a logical-ordering question. Given a
set of pairwise constraints (e.g. Alice is in front of Bob), a model should
answer transitive-closure queries (Is Carol in front of Dave?) consistently
whether the constraints are stated using a relation or its inverse.
Each item exists as a paired (original, equivalent) record… See the full description on the dataset page: https://huggingface.co/datasets/jizej/Competence-Based-Evaluation.NQ-RAG-DPO-Evaluation
Dataset Card
Dataset Summary
This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO).
The system is organized into three interconnected pipelines:
1️. RAG Pipeline
The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark.
For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.big-red-bark-chat-evaluation
Big Red Bark Chat Q&A Dataset
Dataset Description
This dataset contains 12,385 question-and-answer pairs collected from Big Red Bark Chat, an innovative AI assistant developed at Cornell University that answers questions about dog health (as well as other animal species). While it does not replace professional veterinary advice, it serves as a valuable starting point by searching trusted sources. Big Red Bark Chat is designed to provide quick and reliable answers… See the full description on the dataset page: https://huggingface.co/datasets/Sr523/big-red-bark-chat-evaluation.less-is-moe-gpqa-diamond-evaluation
Less-is-MoE GPQA-Diamond evaluation set
This private dataset stores the 198-question GPQA-Diamond evaluation file used
by the MoE-Honing evaluation format.
Upstream source: Idavidrein/gpqa, config gpqa_diamond
Upstream revision: 633f5ee89ab8ad4522a9f850766b73f62147ffdd
Split: test
Rows: 198
SHA-256: d5b0d6dad6c1993a8cb17fd7aa635fbc5e8b684ae42b5b6e80d15467b399eec3
Fields: problem, solution, domain
The problem field contains the formatted four-choice prompt, solution stores
the… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-gpqa-diamond-evaluation.pedants_qa_evaluation_bench
pedants_qa_evaluation
This dataset evaluates candidate answers for various question-answering (QA) tasks across multiple datasets such as Jeopardy!, hotpotQA, nq-open, narrativeQA, and BIOMRC, etc. See details in paper. It contains questions, reference answers (ground truth), model-generated candidate answers, and human judgments indicating whether the candidate answers are correct.
Dataset Details
Column
Type
Description
question
string
The question asked… See the full description on the dataset page: https://huggingface.co/datasets/zli12321/pedants_qa_evaluation_bench.mobile_sft_evaluation
Mobile Sft Evaluation
Dataset Description
Mobile QA evaluation dataset with 200 randomly sampled questions and rewritten answers for SFT model evaluation. Answers maintain semantic equivalence with different phrasing for robust evaluation.
Dataset Summary
Total Examples: 200
Task: Question Answering
Language: English
Format: JSONL (one JSON object per line)
Dataset Structure
Example Entry
{
"question": "How can mobile technology expand… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/mobile_sft_evaluation.
