CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HiTZ /casimedicos-exp Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments for the correct answer but also arguments to explain why the remaining possible answers are incorrect. This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation. The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.tabulartext-generation1K<n<10K4 likes1.9k downloads3y agoHugging Face02HiTZ /MedExpQA MexExpQA: Multilingual Benchmarking of Medical QA with reference gold explanations and Retrieval Augmented Generation (RAG) We present a new multilingual parallel medical benchmark, MedExpQA, for the evaluation of LLMs on Medical Question Answering. This benchmark can be used for various NLP tasks including: Medical Question Answering or Explanation Generation. Although the design of MedExpQA is independent of any specific dataset, for the first version of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MedExpQA.tabulartext-generation1K<n<10K9 likes1.9k downloads2y agoHugging Face03HiTZ /EusExams Dataset Card for EusExams [!WARNING] A newer version of this dataset is available! Please use EusExams-v2 which features deduplication, data grouping, and new data. EusExams is a collection of tests designed to prepare individuals for Public Service examinations conducted by several Basque institutions, including the public health system Osakidetza, the Basque Government, the City Councils of Bilbao and Gasteiz, and the University of the Basque Country (UPV/EHU). Within each… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusExams.textquestion-answering10K<n<100K2 likes656 downloads3mo agoHugging Face04HiTZ /BertaQA Dataset Card for BertaQA BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.tabularquestion-answering10K<n<100K1 likes558 downloads2y agoHugging Face05HiTZ /truthfulqa-multi Dataset Card for TruthfulQA-multi TruthfulQA-multi is a professionally translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages. Dataset Details Dataset Description TruthfulQA-multi extends the original English TruthfulQA dataset to four additional languages… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi.textquestion-answering1K<n<10K2 likes237 downloads1y agoHugging Face06HiTZ /EusProficiency Dataset Card for EusProficiency EusProficiency comprises 5,169 exercises on different topics from past EGA exams, the official C1-level certificate of proficiency in Basque. We collected the atarikoa exercises from EGA exams through the years 1998 to 2008. Atarikoa is the first qualifying test of EGA, which measures different aspects of language competency, such as reading comprehension, grammar, vocabulary, spelling, and writing. Each test generally has 85 multiple-choice questions… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusProficiency.tabularquestion-answering1K<n<10K2 likes215 downloads2y agoHugging Face07HiTZ /EusReading Dataset Card for EusReading EusReading consists of 352 reading comprehension exercises (irakurmena) sourced from the set of past EGA exams from 1998 to 2008. Each test generally has 10 multiple-choice questions, with 4 choices and a single correct answer. These exercises are more challenging than Belebele due to the complexity and length of the input texts. As a result, EusReading is useful to measure long context understanding of models. Curated by: HiTZ Research Center & IXA… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusReading.tabularquestion-answeringn<1K2 likes207 downloads2y agoHugging Face08HiTZ /casimedicos-arg CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures CasiMedicos-Arg is, to the best of our knowledge, the first multilingual dataset for Medical Question Answering where correct and incorrect diagnoses for a clinical case are enriched with a natural language explanation written by doctors. The casimedicos-exp have been manually annotated with argument components (i.e., premise, claim) and argument… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-arg.texttext-generation1K<n<10K1 likes171 downloads3mo agoHugging Face09HiTZ /bbq BBQ Dataset The Bias Benchmark for Question Answering (BBQ) dataset evaluates social biases in language models through question-answering tasks in English. Dataset Description This dataset contains questions designed to test for social biases across multiple demographic dimensions. Each question comes in two variants: Ambiguous (ambig): Questions where the correct answer should be "unknown" due to insufficient information Disambiguated (disambig): Questions with… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/bbq.tabularquestion-answering10K<n<100K0 likes159 downloads1y agoHugging Face10HiTZ /ARC-eu Dataset Card for ARC-eu Point of Contact: hitz@ehu.eus Dataset Description Dataset Summary ARC-eu is the professional translation to Basque of ARC's (Clark et al., 2018) validation and test partitions. ARC is a QA benchmark of grade-school level, multiple-choice science questions. Languages eu-ES Dataset Structure Data Instances ARC-eu examples look like this: { "id": "MCAS_2000_4_6", "question": "Zein teknologia… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ARC-eu.textquestion-answering1K<n<10K0 likes158 downloads2y agoHugging Face11HiTZ /EusTrivia Dataset Card for EusTrivia EusTrivia consists of 1,715 trivia questions from multiple online sources. 56.3% of the questions are elementary level (grades 3-6), while the rest are considered challenging. A significant portion of the questions focus specifically on the Basque Country, its language and culture. Each multiple-choice question contains two, three or four choices (3.84 on average) and a single correct answer. Five areas of knowledge are covered: Humanities and Natural… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusTrivia.tabularquestion-answering1K<n<10K1 likes155 downloads2y agoHugging Face12HiTZ /Multilingual-BioASQ-6B Mutilingual BioASQ-6B We translate the BioASQ-6B English Question Answering dataset to generate parallel French, Italian and Spanish versions using the NLLB200 3B parameter model. For more info read the original task description: [http://bioasq.org/participate/challenges_year_6](http://bioasq.org/participate/challenges_year_6) We translate the body, snippets, ideal_answer and exact_answer fields. We have validated the quality of the ideal_answer field, however, the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Multilingual-BioASQ-6B.textquestion-answering10K<n<100K2 likes130 downloads2y agoHugging Face13HiTZ /casimedicos-squad Antidote CasiMedicos in SQuAD Format for Explanatory Argument Extraction We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments for the correct answer but also arguments to explain why the remaining possible answers are incorrect. Furthermore, this dataset allows us to setup a novel extractive task which consists of identifying the explanation of the correct answer written by medical doctors.… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-squad.question-answering1K<n<10K1 likes105 downloads2y agoHugging Face14HiTZ /PIQA-eu Dataset Card for PIQA-eu Point of Contact: hitz@ehu.eus Dataset Description Dataset Summary PIQA-eu is the professional translation to Basque of the PIQA's (Bisk et al., 2020) validation partition. PIQA is a commonsense QA benchmark for naive physics reasoning focusing on how we interact with everyday objects in everyday situations. Languages eu-ES Dataset Structure Data Instances PIQA-eu examples look like this: {… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/PIQA-eu.tabularquestion-answering1K<n<10K0 likes92 downloads2y agoHugging Face15HiTZ /elkarhizketak-RAG Dataset Card for ElkarHizketak RAG and its Disruptor Variants Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts). Dataset Details Dataset Description This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.tabularquestion-answering1K<n<10K1 likes78 downloads3mo agoHugging Face16HiTZ /EusExams-v2 Dataset Card for EusExams-v2 EusExams-v2 is an updated and refined collection of tests designed to prepare individuals for Public Service examinations conducted by several Basque institutions, including the public health system Osakidetza, the Basque Government, the City Councils of Bilbao and Gasteiz, and the University of the Basque Country (UPV/EHU). Within each of these groups, there are different exams for public positions, such as administrative and assistant roles. Each… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusExams-v2.textquestion-answering10K<n<100K0 likes59 downloads3mo agoHugging Face17HiTZ /IBERtaQAgated Dataset Card for IBERtaQA Curated by: HiTZ Center, University of the Basque Country (EHU) Barcelona Supercomputing Center (BSC) CiTIUS, University of Santiago de Compostela (USC) GPLSI, University of Alicante (UA) Funded by: Project Desarrollo de Modelos ALIA; Projecte AINA License: CC-BY-4.0 Dataset Summary IBERtaQA is a multilingual benchmark designed to evaluate the cultural and factual knowledge of language models across Iberian languages and cultural… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/IBERtaQA.tabularquestion-answering10K<n<100K0 likes50 downloads5d agoHugging Face18HiTZ /elkarhizketak Dataset Card for ElkarHizketak Dataset Summary ElkarHizketak is a low resource conversational Question Answering (QA) dataset in Basque created by Basque speaker volunteers. The dataset contains close to 400 dialogues and more than 1600 question and answers, and its small size presents a realistic low-resource scenario for conversational QA systems. The dataset is built on top of Wikipedia sections about popular people and organizations. The dialogues involve two crowd… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak.textquestion-answering1K<n<10K0 likes41 downloads1y agoHugging Face19ixa-hitz /elkarhizketakElkarHizketak is a low resource conversational Question Answering (QA) dataset in Basque created by Basque speaker volunteers. The dataset contains close to 400 dialogues and more than 1600 question and answers, and its small size presents a realistic low-resource scenario for conversational QA systems. The dataset is built on top of Wikipedia sections about popular people and organizations. The dialogues involve two crowd workers: (1) a student ask questions after reading a small introduction about the person, but without seeing the section text; and (2) a teacher answers the questions selecting a span of text of the section.question-answering1K<n<10K1 likes37 downloads3y agoHugging Face20HiTZ /metaphor-llms metaphorLLM This repository includes the data used in the paper Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding. (ACL Findings, 2025). Code is also available in GitHub. Our paper presents a comprehensive evaluation of the capabilities of Large Language Models (LLMs) in metaphor interpretation across multiple datasets, tasks, and prompt configurations. The results indicate that LLMs' performance is more influenced by features like lexical… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/metaphor-llms.text-classification1 likes36 downloads1y agoHugging Face21HiTZ /TOOLtifrutiThe Basque evaluation ecosystem still lacks standardized datasets and protocols to assess agentic behavior, and in particular tool selection and tool use in end-to-end Agentic RAG settings. To address this gap, we introduce TOOLtifruti, an ad hoc dataset designed to evaluate whether an LLM can identify when a tool is needed and select the appropriate tool among multiple domain-specific options in our use case. This setup makes tool-calling evaluation straightforward and reproducible, and it… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/TOOLtifruti.question-answering1K<n<10K1 likes31 downloads8mo agoHugging Face22HiTZ /truthfulqa-multi-MT Dataset Card for TruthfulQA-multi MT TruthfulQA-multi is an automatically translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages. Dataset Details Dataset Description TruthfulQA-multi extends the original English TruthfulQA dataset to four additional… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi-MT.textquestion-answering1K<n<10K0 likes29 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.