CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HiTZ /casimedicos-exp Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments for the correct answer but also arguments to explain why the remaining possible answers are incorrect. This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation. The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.tabulartext-generation1K<n<10K4 likes1.9k downloads3y agoHugging Face02HiTZ /MedExpQA MexExpQA: Multilingual Benchmarking of Medical QA with reference gold explanations and Retrieval Augmented Generation (RAG) We present a new multilingual parallel medical benchmark, MedExpQA, for the evaluation of LLMs on Medical Question Answering. This benchmark can be used for various NLP tasks including: Medical Question Answering or Explanation Generation. Although the design of MedExpQA is independent of any specific dataset, for the first version of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MedExpQA.tabulartext-generation1K<n<10K9 likes1.9k downloads2y agoHugging Face03HiTZ /Multilingual-Medical-Corpus Mutilingual Medical Corpus Multilingual-Medical-Corpus a 3 billion word multilingual corpus for training LLMs adapted to the medical domain. Multilingual-Medical-Corpus includes four languages, namely, English, Spanish, French, and Italian. 📖 Paper: Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain 🌐 Project Website: https://univ-cotedazur.eu/antidote Corpus Description Developed by: Iker García-Ferrero, Rodrigo Agerri… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Multilingual-Medical-Corpus.text10M<n<100M45 likes1.2k downloads2y agoHugging Face04HiTZ /composite_corpus_es_v1.0 Composite dataset for Spanish made from public available data This dataset is composed of the following public available data: Train split: The train split is composed of the following datasets combined: mozilla-foundation/common_voice_18_0/es: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data) openslr: a train split made from the SLR(39,61,67,71,72,73,74,75,108) subsets… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_es_v1.0.audioautomatic-speech-recognition100K<n<1M0 likes694 downloads1y agoHugging Face05HiTZ /EusExams Dataset Card for EusExams [!WARNING] A newer version of this dataset is available! Please use EusExams-v2 which features deduplication, data grouping, and new data. EusExams is a collection of tests designed to prepare individuals for Public Service examinations conducted by several Basque institutions, including the public health system Osakidetza, the Basque Government, the City Councils of Bilbao and Gasteiz, and the University of the Basque Country (UPV/EHU). Within each… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusExams.textquestion-answering10K<n<100K2 likes656 downloads3mo agoHugging Face06HiTZ /latxa-corpus-v1.1 Latxa Corpus v1.1 This is the training corpus of Latxa v1.1, a family of large language models for Basque based on Llama 2. 💻 Repository: https://github.com/hitz-zentroa/latxa 📒 Blog Post: Latxa: An Open Language Model and Evaluation Suite for Basque 📖 Paper: Latxa: An Open Language Model and Evaluation Suite for Basque 📧 Point of Contact: hitz@ehu.eus 📌 Notice As of February 13th 2026, this repository reflects a curated version of the original dataset. Some data… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v1.1.textfill-mask1M<n<10M2 likes565 downloads8mo agoHugging Face07HiTZ /BertaQA Dataset Card for BertaQA BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.tabularquestion-answering10K<n<100K1 likes558 downloads2y agoHugging Face08HiTZ /composite_corpus_eseu_v1.0 Composite bilingual dataset for Spanish and Basque made from public available data This dataset is composed of the following public available data: Train split: The train split is composed of the following datasets combined: mozilla-foundation/common_voice_18_0/es: a portion of the "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data) mozilla-foundation/common_voice_18_0/eu:… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eseu_v1.0.audioautomatic-speech-recognition100K<n<1M2 likes536 downloads1y agoHugging Face09HiTZ /Magpie-Llama-3.1-70B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-70B-Instruc with the MAGPIE codebase. The filtered dataset can be found here: HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered System prompts used General <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n Code <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-70B-Instruct-Unfiltered.tabular1M<n<10M0 likes410 downloads1y agoHugging Face10HiTZ /Magpie-Llama-3.1-8B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-8B-Instruc with the MAGPIE codebase. The filtered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered System prompts used General <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n Code <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered.tabular1M<n<10M0 likes398 downloads1y agoHugging Face11HiTZ /magpie-en-eu-reasoning-instructions-qwen3 Dataset Card for magpie-en-eu-reasoning-instructions-qwen3 Dataset Summary The magpie-en-eu-reasoning-instructions-qwen3 dataset is a large-scale, high-quality, bilingual instruction and preference dataset developed by the HiTZ Center. It is specifically tailored for training, aligning, and evaluating reasoning-focused Large Language Models (LLMs) in both English and Basque (Euskera). Built using the self-synthesizing Magpie methodology, the dataset contains a… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/magpie-en-eu-reasoning-instructions-qwen3.text1M<n<10M1 likes377 downloads3mo agoHugging Face12HiTZ /composite_corpus_eu_v2.1 Composite dataset for Basque made from public available data This dataset is composed of the following public available data: Train split: The train split is composed of the following datasets combined: mozilla-foundation/common_voice_18_0/eu: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data) gttsehu/basque_parliament_1/eu: "train_clean" split removing some of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eu_v2.1.audioautomatic-speech-recognition100K<n<1M3 likes314 downloads2y agoHugging Face13HiTZ /latxa-corpus-v2 Latxa Corpus v2 📧 Point of Contact: hitz@ehu.eus Dataset Summary Curated by: HiTZ Research Center & IXA Research group (University of the Basque Country UPV/EHU) Language(s): eu-ES Latxa Corpus v2 is a large-scale monolingual Basque corpus, created by combining curated crawls, public datasets, institutional data, and newly collected resources. Compared to v1.1, it substantially increases coverage, diversity, and volume. The final corpus is deduplicated, filtered, and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v2.textfill-mask1M<n<10M1 likes276 downloads8mo agoHugging Face14HiTZ /xnli-eu Dataset Card for XNLIeu XNLIeu is an extension of XNLI translated from English to Basque. It has been designed as a cross-lingual dataset for the Natural Language Inference task, a text-classification task that consists on classifying pairs of sentences, a premise and a hypothesis, according to their semantic relation out of three possible labels: entailment, contradiction and neutral. Dataset Details Dataset Description XNLI is a popular Natural… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/xnli-eu.text100K<n<1M0 likes271 downloads6mo agoHugging Face15HiTZ /truthfulqa-multi Dataset Card for TruthfulQA-multi TruthfulQA-multi is a professionally translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages. Dataset Details Dataset Description TruthfulQA-multi extends the original English TruthfulQA dataset to four additional languages… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi.textquestion-answering1K<n<10K2 likes237 downloads1y agoHugging Face16HiTZ /Magpie-Llama-3-70B-Instruct-UnfilteredDataset generated using meta-llama/Meta-Llama-3-70B-Instruct with the MAGPIE codebase. The filtered dataset can be found here: HiTZ/Magpie-Llama-3-70B-Instruct-Filtered System prompts used General <|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n Code <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI assistant designed to provide helpful, step-by-step guidance on coding problems. The user will ask you a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3-70B-Instruct-Unfiltered.tabular1M<n<10M0 likes224 downloads2y agoHugging Face17HiTZ /EusProficiency Dataset Card for EusProficiency EusProficiency comprises 5,169 exercises on different topics from past EGA exams, the official C1-level certificate of proficiency in Basque. We collected the atarikoa exercises from EGA exams through the years 1998 to 2008. Atarikoa is the first qualifying test of EGA, which measures different aspects of language competency, such as reading comprehension, grammar, vocabulary, spelling, and writing. Each test generally has 85 multiple-choice questions… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusProficiency.tabularquestion-answering1K<n<10K2 likes215 downloads2y agoHugging Face18HiTZ /EusReading Dataset Card for EusReading EusReading consists of 352 reading comprehension exercises (irakurmena) sourced from the set of past EGA exams from 1998 to 2008. Each test generally has 10 multiple-choice questions, with 4 choices and a single correct answer. These exercises are more challenging than Belebele due to the complexity and length of the input texts. As a result, EusReading is useful to measure long context understanding of models. Curated by: HiTZ Research Center & IXA… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusReading.tabularquestion-answeringn<1K2 likes207 downloads2y agoHugging Face19HiTZ /multilingual-abstrct Mutilingual AbstRCT We translate the AbstRCT English Argument Mining Dataset dataset to generate parallel French, Italian and Spanish versions using the NLLB200 3B parameter model and projected using word alignment tools. The projections have been manually corrected. For more info about the original English AbstRCT dataset read the original paper. For the translation and projection data see https://github.com/ragerri/abstrct-projections/tree/final. 📖 Paper: Medical… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/multilingual-abstrct.texttoken-classification10K<n<100K0 likes205 downloads2y agoHugging Face20open-llm-leaderboard-old /details_HiTZ__alpaca-lora-65b-en-pt-es-ca Dataset Card for Evaluation run of HiTZ/alpaca-lora-65b-en-pt-es-ca Dataset Summary Dataset automatically created during the evaluation run of model HiTZ/alpaca-lora-65b-en-pt-es-ca on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_HiTZ__alpaca-lora-65b-en-pt-es-ca.0 likes200 downloads3y agoHugging Face21HiTZ /MATE Dataset Card for HiTZ/MATE This dataset provides a benchmark consisting of 5,500 question-answering examples to assess the cross-modal entity linking capabilities of vision-language models (VLMs). The ability to link entities in different modalities is measured in a question-answering setting, where each scene is represented in both the visual modality (image) and the textual one (a list of objects and their attributes in JSON format). The dataset is provided in two configurations:… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MATE.imagevisual-question-answering10K<n<100K0 likes183 downloads1y agoHugging Face22HiTZ /casimedicos-arg CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures CasiMedicos-Arg is, to the best of our knowledge, the first multilingual dataset for Medical Question Answering where correct and incorrect diagnoses for a clinical case are enriched with a natural language explanation written by doctors. The casimedicos-exp have been manually annotated with argument components (i.e., premise, claim) and argument… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-arg.texttext-generation1K<n<10K1 likes171 downloads3mo agoHugging Face23HiTZ /bbq BBQ Dataset The Bias Benchmark for Question Answering (BBQ) dataset evaluates social biases in language models through question-answering tasks in English. Dataset Description This dataset contains questions designed to test for social biases across multiple demographic dimensions. Each question comes in two variants: Ambiguous (ambig): Questions where the correct answer should be "unknown" due to insufficient information Disambiguated (disambig): Questions with… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/bbq.tabularquestion-answering10K<n<100K0 likes159 downloads1y agoHugging Face24HiTZ /ARC-eu Dataset Card for ARC-eu Point of Contact: hitz@ehu.eus Dataset Description Dataset Summary ARC-eu is the professional translation to Basque of ARC's (Clark et al., 2018) validation and test partitions. ARC is a QA benchmark of grade-school level, multiple-choice science questions. Languages eu-ES Dataset Structure Data Instances ARC-eu examples look like this: { "id": "MCAS_2000_4_6", "question": "Zein teknologia… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ARC-eu.textquestion-answering1K<n<10K0 likes158 downloads2y agoHugging Face25HiTZ /Magpie-Llama-3-70B-Instruct-FilteredDataset generated using meta-llama/Meta-Llama-3-70B-Instruct with the MAGPIE codebase. The unfiltered dataset can be found here: HiTZ/Magpie-Llama-3-70B-Instruct-Unfiltered Filter criteria def high_quality_filter(example): return ( example["input_quality"] in ["good", "excellent", "average"] and example["instruct_reward"] > -10 and not example["instruction"].endswith(":") and ( example["min_similar_conversation_id"] is None… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3-70B-Instruct-Filtered.tabular1M<n<10M0 likes158 downloads2y agoHugging Face26HiTZ /EusTrivia Dataset Card for EusTrivia EusTrivia consists of 1,715 trivia questions from multiple online sources. 56.3% of the questions are elementary level (grades 3-6), while the rest are considered challenging. A significant portion of the questions focus specifically on the Basque Country, its language and culture. Each multiple-choice question contains two, three or four choices (3.84 on average) and a single correct answer. Five areas of knowledge are covered: Humanities and Natural… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusTrivia.tabularquestion-answering1K<n<10K1 likes155 downloads2y agoHugging Face27HiTZ /MGSM-eu Dataset Card for MGSM-eu Point of Contact: hitz@ehu.eus Dataset Description Dataset Summary MGSM (Shi et al., 2023) is a subset of 250 grade-school math problems from the GSM8K dataset (Cobbe et al., 2021) that has been manually translated into 10 typologically diverse languages. Here, we provide professional translations to yet another language: Basque. Languages eu-ES Dataset Structure Data Instances MGSM-eu train examples… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MGSM-eu.textn<1K0 likes151 downloads2y agoHugging Face28HiTZ /Multilingual-BioASQ-6B Mutilingual BioASQ-6B We translate the BioASQ-6B English Question Answering dataset to generate parallel French, Italian and Spanish versions using the NLLB200 3B parameter model. For more info read the original task description: [http://bioasq.org/participate/challenges_year_6](http://bioasq.org/participate/challenges_year_6) We translate the body, snippets, ideal_answer and exact_answer fields. We have validated the quality of the ideal_answer field, however, the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Multilingual-BioASQ-6B.textquestion-answering10K<n<100K2 likes130 downloads2y agoHugging Face29HiTZ /meta4xnli Dataset Card for Dataset Name Meta4XNLI is a parallel dataset with annotations in English and Spanish for metaphor detection at token level (13320 sentences) and metaphor interpretation framed within NLI the task (9990 premise-hypothesis pairs). It is a collection of existing NLI datasets manually labeled for both metaphor tasks. Repository: data available also in .tsv format at https://github.com/elisanchez-beep/meta4xnli Paper: Meta4XNLI: A Crosslingual Parallel Corpus for… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/meta4xnli.texttoken-classification10K<n<100K1 likes126 downloads2y agoHugging Face30HiTZ /CROQ 🌍🏺 CROQ: Culture-Related Open Questions A multilingual benchmark for evaluating cultural and regional biases in large language models through open-ended cultural questions. CROQ (Culture-Related Open Questions) is a multilingual dataset designed to uncover cultural and regional biases in large language models (LLMs). Unlike traditional cultural benchmarks based on multiple-choice or factual questions, CROQ focuses on open-ended cultural questions that have no single correct… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CROQ.texttext-generation1K<n<10K1 likes121 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.