datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
casimedicos-arg
CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures
CasiMedicos-Arg is, to the best of our knowledge, the first
multilingual dataset for Medical Question Answering where correct and incorrect diagnoses for a clinical case are
enriched with a natural language explanation written by doctors.
The casimedicos-exp have been manually annotated with
argument components (i.e., premise, claim) and argument… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-arg.CROQ
🌍🏺 CROQ: Culture-Related Open Questions
A multilingual benchmark for evaluating cultural and regional biases in large language models through open-ended cultural questions.
CROQ (Culture-Related Open Questions) is a multilingual dataset designed to uncover cultural and regional biases in large language models (LLMs). Unlike traditional cultural benchmarks based on multiple-choice or factual questions, CROQ focuses on open-ended cultural questions that have no single correct… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CROQ.CONAN-EUSContent Warning: This dataset contains examples of offensive language that do not reflect the authors’ views
CONAN-EUS: Basque and Spanish Parallel Counter Narratives Dataset
CONAN-EUS was created by professionally translating all 6654 English HS-CN pairs of the original CONAN dataset into
Basque and Spanish. For experimentation we generated train, validation and test splits in a way that no HS-CN pairs occurred across them.
CONAN-EUS Splits
Total HS-CN… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CONAN-EUS.basqueparl
BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions
This repository contains BasqueParl, a bilingual corpus for political discourse analysis. It covers transcriptions from the Parliament of
the Basque Autonomous Community for eight years and two legislative terms (2012-2020), and its main characteristic is the presence of Basque-Spanish
code-switching speeches.
📖 Paper: BasqueParl A Bilingual Corpus of Basque Parliamentary Transcriptions In LREC 2022.… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/basqueparl.
