datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
casimedicos-exp
Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams
We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments
for the correct answer but also arguments to explain why the remaining possible answers are incorrect.
This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation.
The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.MedExpQA
MexExpQA: Multilingual Benchmarking of Medical QA with reference gold explanations and Retrieval Augmented Generation (RAG)
We present a new multilingual parallel medical benchmark, MedExpQA, for the evaluation of LLMs on Medical Question Answering.
This benchmark can be used for various NLP tasks including: Medical Question Answering or Explanation Generation.
Although the design of MedExpQA is independent of any specific dataset, for the first version of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MedExpQA.cvefixes
CVEfixes Security Vulnerabilities Dataset
Security vulnerability data from CVEfixes v1.0.8 with 12,987 vulnerability fix records across 11,726 unique CVEs and 4,205 repositories.
Contains CVE metadata (descriptions, CVSS scores, CWE classifications), git commit data, and code diffs showing vulnerable vs fixed code.
Usage
from datasets import load_dataset
dataset = load_dataset("hitoshura25/cvefixes")
Citation
If you use this dataset, please cite the original… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/cvefixes.latxa-corpus-v1.1
Latxa Corpus v1.1
This is the training corpus of Latxa v1.1, a family of large language models for Basque based on Llama 2.
💻 Repository: https://github.com/hitz-zentroa/latxa
📒 Blog Post: Latxa: An Open Language Model and Evaluation Suite for Basque
📖 Paper: Latxa: An Open Language Model and Evaluation Suite for Basque
📧 Point of Contact: hitz@ehu.eus
📌 Notice
As of February 13th 2026, this repository reflects a curated version of the original dataset.
Some data… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v1.1.latxa-corpus-v2
Latxa Corpus v2
📧 Point of Contact: hitz@ehu.eus
Dataset Summary
Curated by: HiTZ Research Center & IXA Research group (University of the Basque Country UPV/EHU)
Language(s): eu-ES
Latxa Corpus v2 is a large-scale monolingual Basque corpus, created by combining curated crawls, public datasets, institutional data, and newly collected resources.
Compared to v1.1, it substantially increases coverage, diversity, and volume.
The final corpus is deduplicated, filtered, and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v2.CROQ
🌍🏺 CROQ: Culture-Related Open Questions
A multilingual benchmark for evaluating cultural and regional biases in large language models through open-ended cultural questions.
CROQ (Culture-Related Open Questions) is a multilingual dataset designed to uncover cultural and regional biases in large language models (LLMs). Unlike traditional cultural benchmarks based on multiple-choice or factual questions, CROQ focuses on open-ended cultural questions that have no single correct… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CROQ.crossvul
CrossVul Multi-Language Security Vulnerability Dataset
Security vulnerability dataset from CrossVul with 9,313 before/after code pairs across 158 CWE categories and 21 programming languages.
Contains vulnerable code examples paired with their secure fixes, ideal for training AI models on security code remediation.
Dataset Statistics
Total Examples: 9,313
CWE Categories: 158
Languages: 21
Format: Raw vulnerability records (JSON Lines)
Top Languages… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/crossvul.casimedicos-arg
CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures
CasiMedicos-Arg is, to the best of our knowledge, the first
multilingual dataset for Medical Question Answering where correct and incorrect diagnoses for a clinical case are
enriched with a natural language explanation written by doctors.
The casimedicos-exp have been manually annotated with
argument components (i.e., premise, claim) and argument… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-arg.megavul
MegaVul Dataset (CVEfixes-Compatible Format)
Dataset Description
This is a processed version of the MegaVul dataset converted to match the CVEfixes schema format for unified vulnerability analysis and model training.
Source Kaggle Dataset: marcdamie/megavul-a-cc-java-vulnerability-dataset
Original Project: Icyrockton/MegaVul
Dataset Summary
MegaVul is a large, high-quality, extensible, continuously updated C/C++/Java function-level vulnerability dataset.… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/megavul.hitchcock-psycho-1960-film-dataset-transformed
Psycho → AI-Model Dataset (Transformed)
A thematic re-skin of the Psycho (1960) Q&A dataset into an original AI-model setting where the world is transformed into an AI/data-center environment.
Character names, actor names, objects, locations, production references, dates, and thematic elements are remapped to AI/ML concepts and modern technology.
File: psycho_dataset_transformed.jsonl
Format: JSONL — one JSON object per line
Schema: each line has prompt and completion string… See the full description on the dataset page: https://huggingface.co/datasets/antfr99/hitchcock-psycho-1960-film-dataset-transformed.Polaris-Hard
Polaris-Hard
A random subset of Polaris-Dataset-53K, with 3,200 distinct math problems in a single train split.
Original difficulty
Source pool
Eligible pool
Selected
Share
0/8
15,368
9,331
2,000
62.5%
1/8
6,956
4,929
1,200
37.5%
Total
22,324
14,260
3,200
100%
Sampling
Source revision: 296f8e34132e63f4a1d70e0dcc8bddebb43f03e4.
Seed: 42. Uniform random sampling without replacement within each group, followed by a deterministic shuffle of the… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Polaris-Hard.CONAN-EUSContent Warning: This dataset contains examples of offensive language that do not reflect the authors’ views
CONAN-EUS: Basque and Spanish Parallel Counter Narratives Dataset
CONAN-EUS was created by professionally translating all 6654 English HS-CN pairs of the original CONAN dataset into
Basque and Spanish. For experimentation we generated train, validation and test splits in a way that no HS-CN pairs occurred across them.
CONAN-EUS Splits
Total HS-CN… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CONAN-EUS.Polaris-1-8-3200
Polaris-1-8-3200
A random subset of Polaris-Dataset-53K, with 3,200 distinct math problems in a single train split.
Original difficulty
Source pool
Eligible pool
Selected
Share
1/8
6,956
4,929
3,200
100.0%
Total
6,956
4,929
3,200
100%
Sampling
Source revision: 296f8e34132e63f4a1d70e0dcc8bddebb43f03e4.
Seed: 42. Uniform random sampling without replacement within each group, followed by a deterministic shuffle of the combined selection.… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Polaris-1-8-3200.euscrawlEusCrawl (http://www.ixa.eus/euscrawl/) is a high-quality corpus for
Basque comprising 12.5 million documents and 423 million tokens,
totalling 2.1 GiB of uncompressed text. EusCrawl was built using
ad-hoc scrapers to extract text from 33 Basque websites with
high-quality content, resulting in cleaner text compared to general
purpose approaches.
We do not claim ownership of any document in the corpus. All documents
we collected were published under a Creative Commons license in their
original website, and the specific variant can be found in the
"license" field of each document. Should you consider
that our data contains material that is owned by you and you would not
like to be reproduced here, please contact Aitor Soroa at
a.soroa@ehu.eus.
For more details about the corpus, refer to our paper "Artetxe M.,
Aldabe I., Agerri R., Perez-de-Viñaspre O, Soroa A. (2022). Does
Corpus Quality Really Matter for Low-Resource Languages?"
https://arxiv.org/abs/2203.08111
If you use our corpus or models for academic research, please cite the paper in question:
@misc{artetxe2022euscrawl,
title={Does corpus quality really matter for low-resource languages?},
author={Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de-Viñaspre, Aitor Soroa},
year={2022},
eprint={2203.08111},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
For questions please contact Aitor Soroa at a.soroa@ehu.eus.BASSE
BASSE: BAsque and Spanish Summarization Evaluation
BASSE is a multilingual (Basque and Spanish) dataset designed primarily for the
meta-evaluation of automatic summarization metrics and LLM-as-a-Judge models.
Dataset Details
Dataset Description
BASSE is a multilingual (Basque and Spanish) dataset designed primarily for the
meta-evaluation of automatic summarization metrics and LLM-as-a-Judge models.
We generated automatic summaries for 90 news documents in… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BASSE.elkarhizketak-RAG
Dataset Card for ElkarHizketak RAG and its Disruptor Variants
Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts).
Dataset Details
Dataset Description
This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.alpaca_mtAlpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better. This dataset also includes machine-translated data for 6 Iberian languages: Portuguese, Spanish, Catalan, Basque, Galician and Asturian.BasqueSumm
BasqueSumm
BasqueSumm was automatically compiled from www.berria.eus
using trafilatura to extract the texts.
Each instance has the following key-value pairs:
"date" (str): When the article was published, formatted as "yyyy-mm-dd".
"url" (str): The URL of the original publication.
"category" (str): the articles topic, e.g., economy, society.
"title" (str): The title of the article.
"subtitle" (str): The subtitle of the article.
"summary" (str): The combined title + subtitle, which… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BasqueSumm.Polaris-Test
Polaris-Test
100 randomly sampled 1/8 Polaris questions, held out from
hi-todayis-jh/Polaris-1-8-3200.
The single split is test.
The source is POLARIS-Project/Polaris-Dataset-53K.
This test set uses the same finite-real-no-explicit-proof-v1 numerical-answer
and explicit-proof-wording filter as the pinned training set. The accepted source
indices are taken directly from that training set's sampling manifest.
After excluding training source IDs, matching question text (Unicode… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Polaris-Test.basqueparl
BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions
This repository contains BasqueParl, a bilingual corpus for political discourse analysis. It covers transcriptions from the Parliament of
the Basque Autonomous Community for eight years and two legislative terms (2012-2020), and its main characteristic is the presence of Basque-Spanish
code-switching speeches.
📖 Paper: BasqueParl A Bilingual Corpus of Basque Parliamentary Transcriptions In LREC 2022.… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/basqueparl.CQs-Gen
Critical Questions Generation Dataset: CQs-Gen
This dataset is designed to benchmark the ability of language models to generate critical questions (CQs) for argumentative texts. Each instance consists of a naturally occurring argumentative intervention paired with multiple reference questions, annotated for their usefulness in challenging the arguments.
Dataset Overview
Number of interventions: 220
Average intervention length: 738.4 characters
Average number of… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CQs-Gen.BERnaT-Diverse
BERnaT: Basque Encoders for Representing Natural Textual Diversity
Submitted to LREC 2026
Abstract
Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally
exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this
paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal,
historical, informal, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BERnaT-Diverse.webauthn-security-training-data-20251014_151917
WebAuthn Security Training Data
High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation.
Dataset Description
This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models.
Format: MLX Chat Messages
This dataset uses the MLX LoRA chat format with explicit role separation:
{
"messages": [
{
"role": "system",
"content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251014_151917.Polaris-Test-4-8
Polaris-Test-4-8
100 questions sampled uniformly without replacement from
edbeeching/Polaris-Dataset-53K-4-8.
The pinned source contains 20,371 rows, with difficulty labels 4/8 through 7/8.
The test split preserves the original problem, answer, and difficulty;
source_index is the zero-based position in the pinned source Parquet.
Selection uses Python random.Random(42).sample(range(20371), 100), retaining
the returned order. There is no answer/difficulty filter and no selection… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Polaris-Test-4-8.ifeval_gl
IFEval GL
Dataset Summary
IFEval GL is a Galician instruction-following evaluation dataset in JSONL format.It contains prompts together with the corresponding instruction identifiers and argument constraints used for evaluation.
Dataset Structure
Split
Rows
Features
train
541
4
Features
Feature
Type
Description
key
integer
Unique example identifier
prompt
string
Instruction prompt in Galician… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ifeval_gl.ifeval_eu
IFEval EU
Dataset Summary
IFEval EU is a Basque instruction-following evaluation dataset in JSONL format.It contains prompts together with the corresponding instruction identifiers and argument constraints used for evaluation.
Dataset Structure
Split
Rows
Features
train
541
4
Features
Feature
Type
Description
key
integer
Unique example identifier
prompt
string
Instruction prompt in Basque… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ifeval_eu.webauthn-security-training-data-20251009_152808
WebAuthn Security Training Data
High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation.
Dataset Description
This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models.
Format: MLX Chat Messages
This dataset uses the MLX LoRA chat format with explicit role separation:
{
"messages": [
{
"role": "system",
"content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251009_152808.TopoTalk-Bench
Network Topology Benchmark Dataset
📊 Dataset Overview
This dataset is designed to test large language models' ability to process network topologies, including two main tasks: building from scratch and modifying existing topologies.
Directory
File Count
Description
origin/
90
Original network topologies (empty + original)
nl/
792
Natural language descriptions (396 Chinese + 396 English)
netjson/
396
Ground Truth NetJSON files
Total: 1278 files… See the full description on the dataset page: https://huggingface.co/datasets/smy-hit/TopoTalk-Bench.dataclaw-peteromallet
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value
Sessions
549… See the full description on the dataset page: https://huggingface.co/datasets/hitlabstudios/dataclaw-peteromallet.film_qa_pairs_dataset
