datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
casimedicos-exp
Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams
We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments
for the correct answer but also arguments to explain why the remaining possible answers are incorrect.
This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation.
The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.MedExpQA
MexExpQA: Multilingual Benchmarking of Medical QA with reference gold explanations and Retrieval Augmented Generation (RAG)
We present a new multilingual parallel medical benchmark, MedExpQA, for the evaluation of LLMs on Medical Question Answering.
This benchmark can be used for various NLP tasks including: Medical Question Answering or Explanation Generation.
Although the design of MedExpQA is independent of any specific dataset, for the first version of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MedExpQA.latxa-corpus-v1.1
Latxa Corpus v1.1
This is the training corpus of Latxa v1.1, a family of large language models for Basque based on Llama 2.
💻 Repository: https://github.com/hitz-zentroa/latxa
📒 Blog Post: Latxa: An Open Language Model and Evaluation Suite for Basque
📖 Paper: Latxa: An Open Language Model and Evaluation Suite for Basque
📧 Point of Contact: hitz@ehu.eus
📌 Notice
As of February 13th 2026, this repository reflects a curated version of the original dataset.
Some data… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v1.1.latxa-corpus-v2
Latxa Corpus v2
📧 Point of Contact: hitz@ehu.eus
Dataset Summary
Curated by: HiTZ Research Center & IXA Research group (University of the Basque Country UPV/EHU)
Language(s): eu-ES
Latxa Corpus v2 is a large-scale monolingual Basque corpus, created by combining curated crawls, public datasets, institutional data, and newly collected resources.
Compared to v1.1, it substantially increases coverage, diversity, and volume.
The final corpus is deduplicated, filtered, and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v2.hitchcock-psycho-1960-film-dataset-transformed
Psycho → AI-Model Dataset (Transformed)
A thematic re-skin of the Psycho (1960) Q&A dataset into an original AI-model setting where the world is transformed into an AI/data-center environment.
Character names, actor names, objects, locations, production references, dates, and thematic elements are remapped to AI/ML concepts and modern technology.
File: psycho_dataset_transformed.jsonl
Format: JSONL — one JSON object per line
Schema: each line has prompt and completion string… See the full description on the dataset page: https://huggingface.co/datasets/antfr99/hitchcock-psycho-1960-film-dataset-transformed.elkarhizketak-RAG
Dataset Card for ElkarHizketak RAG and its Disruptor Variants
Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts).
Dataset Details
Dataset Description
This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.BasqueSumm
BasqueSumm
BasqueSumm was automatically compiled from www.berria.eus
using trafilatura to extract the texts.
Each instance has the following key-value pairs:
"date" (str): When the article was published, formatted as "yyyy-mm-dd".
"url" (str): The URL of the original publication.
"category" (str): the articles topic, e.g., economy, society.
"title" (str): The title of the article.
"subtitle" (str): The subtitle of the article.
"summary" (str): The combined title + subtitle… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BasqueSumm.CQs-Gen
Critical Questions Generation Dataset: CQs-Gen
This dataset is designed to benchmark the ability of language models to generate critical questions (CQs) for argumentative texts. Each instance consists of a naturally occurring argumentative intervention paired with multiple reference questions, annotated for their usefulness in challenging the arguments.
Dataset Overview
Number of interventions: 220
Average intervention length: 738.4 characters
Average number of… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CQs-Gen.BERnaT-Diverse
BERnaT: Basque Encoders for Representing Natural Textual Diversity
Submitted to LREC 2026
Abstract
Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally
exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this
paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal,
historical, informal, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BERnaT-Diverse.ifeval_gl
IFEval GL
Dataset Summary
IFEval GL is a Galician instruction-following evaluation dataset in JSONL format.It contains prompts together with the corresponding instruction identifiers and argument constraints used for evaluation.
Dataset Structure
Split
Rows
Features
train
541
4
Features
Feature
Type
Description
key
integer
Unique example identifier
prompt
string
Instruction prompt in Galician… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ifeval_gl.ifeval_eu
IFEval EU
Dataset Summary
IFEval EU is a Basque instruction-following evaluation dataset in JSONL format.It contains prompts together with the corresponding instruction identifiers and argument constraints used for evaluation.
Dataset Structure
Split
Rows
Features
train
541
4
Features
Feature
Type
Description
key
integer
Unique example identifier
prompt
string
Instruction prompt in Basque… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ifeval_eu.webauthn-security-training-data-20251014_151917
WebAuthn Security Training Data
High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation.
Dataset Description
This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models.
Format: MLX Chat Messages
This dataset uses the MLX LoRA chat format with explicit role separation:
{
"messages": [
{
"role": "system",
"content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251014_151917.webauthn-security-training-data-20251009_152808
WebAuthn Security Training Data
High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation.
Dataset Description
This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models.
Format: MLX Chat Messages
This dataset uses the MLX LoRA chat format with explicit role separation:
{
"messages": [
{
"role": "system",
"content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251009_152808.dataclaw-peteromallet
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value
Sessions
549… See the full description on the dataset page: https://huggingface.co/datasets/hitlabstudios/dataclaw-peteromallet.
