datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
casimedicos-exp
Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams
We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments
for the correct answer but also arguments to explain why the remaining possible answers are incorrect.
This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation.
The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.MedExpQA
MexExpQA: Multilingual Benchmarking of Medical QA with reference gold explanations and Retrieval Augmented Generation (RAG)
We present a new multilingual parallel medical benchmark, MedExpQA, for the evaluation of LLMs on Medical Question Answering.
This benchmark can be used for various NLP tasks including: Medical Question Answering or Explanation Generation.
Although the design of MedExpQA is independent of any specific dataset, for the first version of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MedExpQA.EusExams
Dataset Card for EusExams
[!WARNING]
A newer version of this dataset is available! Please use EusExams-v2 which features deduplication, data grouping, and new data.
EusExams is a collection of tests designed to prepare individuals for Public Service examinations conducted by several Basque institutions, including the public health system Osakidetza, the Basque Government, the City Councils of Bilbao and Gasteiz, and the University of the Basque Country (UPV/EHU). Within each… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusExams.BertaQA
Dataset Card for BertaQA
BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.truthfulqa-multi
Dataset Card for TruthfulQA-multi
TruthfulQA-multi is a professionally translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages.
Dataset Details
Dataset Description
TruthfulQA-multi extends the original English TruthfulQA dataset to four additional languages… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi.EusProficiency
Dataset Card for EusProficiency
EusProficiency comprises 5,169 exercises on different topics from past EGA exams, the official C1-level certificate of proficiency in Basque.
We collected the atarikoa exercises from EGA exams through the years 1998 to 2008. Atarikoa is the first qualifying test of EGA, which measures different aspects of language competency, such as reading comprehension, grammar, vocabulary, spelling, and writing. Each test generally has 85 multiple-choice questions… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusProficiency.EusReading
Dataset Card for EusReading
EusReading consists of 352 reading comprehension exercises (irakurmena) sourced from the set of past EGA exams from 1998 to 2008. Each test generally has 10 multiple-choice questions, with 4 choices and a single correct answer. These exercises are more challenging than Belebele due to the complexity and length of the input texts. As a result, EusReading is useful to measure long context understanding of models.
Curated by: HiTZ Research Center & IXA… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusReading.ARC-eu
Dataset Card for ARC-eu
Point of Contact: hitz@ehu.eus
Dataset Description
Dataset Summary
ARC-eu is the professional translation to Basque of ARC's
(Clark et al., 2018) validation and test partitions.
ARC is a QA benchmark of grade-school level, multiple-choice science questions.
Languages
eu-ES
Dataset Structure
Data Instances
ARC-eu examples look like this:
{
"id": "MCAS_2000_4_6",
"question": "Zein teknologia… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ARC-eu.EusTrivia
Dataset Card for EusTrivia
EusTrivia consists of 1,715 trivia questions from multiple online sources. 56.3% of the questions are elementary level (grades 3-6), while the rest are considered challenging. A significant portion of the questions focus specifically on the Basque Country, its language and culture. Each multiple-choice question contains two, three or four choices (3.84 on average) and a single correct answer. Five areas of knowledge are covered:
Humanities and Natural… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusTrivia.PIQA-eu
Dataset Card for PIQA-eu
Point of Contact: hitz@ehu.eus
Dataset Description
Dataset Summary
PIQA-eu is the professional translation to Basque of the PIQA's
(Bisk et al., 2020) validation partition.
PIQA is a commonsense QA benchmark for naive physics reasoning focusing on how we interact with everyday
objects in everyday situations.
Languages
eu-ES
Dataset Structure
Data Instances
PIQA-eu examples look like this:
{… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/PIQA-eu.elkarhizketak-RAG
Dataset Card for ElkarHizketak RAG and its Disruptor Variants
Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts).
Dataset Details
Dataset Description
This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.EusExams-v2
Dataset Card for EusExams-v2
EusExams-v2 is an updated and refined collection of tests designed to prepare individuals for Public Service examinations conducted by several Basque institutions, including the public health system Osakidetza, the Basque Government, the City Councils of Bilbao and Gasteiz, and the University of the Basque Country (UPV/EHU). Within each of these groups, there are different exams for public positions, such as administrative and assistant roles. Each… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusExams-v2.IBERtaQA
Dataset Card for IBERtaQA
Curated by:
HiTZ Center, University of the Basque Country (EHU)
Barcelona Supercomputing Center (BSC)
CiTIUS, University of Santiago de Compostela (USC)
GPLSI, University of Alicante (UA)
Funded by: Project Desarrollo de Modelos ALIA; Projecte AINA
License: CC-BY-4.0
Dataset Summary
IBERtaQA is a multilingual benchmark designed to evaluate the cultural and factual knowledge of language models across Iberian languages
and cultural… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/IBERtaQA.truthfulqa-multi-MT
Dataset Card for TruthfulQA-multi MT
TruthfulQA-multi is an automatically translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages.
Dataset Details
Dataset Description
TruthfulQA-multi extends the original English TruthfulQA dataset to four additional… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi-MT.
