CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HiTZ /casimedicos-exp Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments for the correct answer but also arguments to explain why the remaining possible answers are incorrect. This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation. The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.tabulartext-generation1K<n<10K4 likes1.9k downloads3y agoHugging Face02HiTZ /MedExpQA MexExpQA: Multilingual Benchmarking of Medical QA with reference gold explanations and Retrieval Augmented Generation (RAG) We present a new multilingual parallel medical benchmark, MedExpQA, for the evaluation of LLMs on Medical Question Answering. This benchmark can be used for various NLP tasks including: Medical Question Answering or Explanation Generation. Although the design of MedExpQA is independent of any specific dataset, for the first version of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MedExpQA.tabulartext-generation1K<n<10K9 likes1.8k downloads2y agoHugging Face03hitoshura25 /cvefixes CVEfixes Security Vulnerabilities Dataset Security vulnerability data from CVEfixes v1.0.8 with 12,987 vulnerability fix records across 11,726 unique CVEs and 4,205 repositories. Contains CVE metadata (descriptions, CVSS scores, CWE classifications), git commit data, and code diffs showing vulnerable vs fixed code. Usage from datasets import load_dataset dataset = load_dataset("hitoshura25/cvefixes") Citation If you use this dataset, please cite the original… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/cvefixes.tabulartext-generation10K<n<100K5 likes1.4k downloads11mo agoHugging Face04HiTZ /latxa-corpus-v1.1 Latxa Corpus v1.1 This is the training corpus of Latxa v1.1, a family of large language models for Basque based on Llama 2. 💻 Repository: https://github.com/hitz-zentroa/latxa 📒 Blog Post: Latxa: An Open Language Model and Evaluation Suite for Basque 📖 Paper: Latxa: An Open Language Model and Evaluation Suite for Basque 📧 Point of Contact: hitz@ehu.eus 📌 Notice As of February 13th 2026, this repository reflects a curated version of the original dataset. Some data… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v1.1.textfill-mask1M<n<10M2 likes558 downloads7mo agoHugging Face05HiTZ /latxa-corpus-v2 Latxa Corpus v2 📧 Point of Contact: hitz@ehu.eus Dataset Summary Curated by: HiTZ Research Center & IXA Research group (University of the Basque Country UPV/EHU) Language(s): eu-ES Latxa Corpus v2 is a large-scale monolingual Basque corpus, created by combining curated crawls, public datasets, institutional data, and newly collected resources. Compared to v1.1, it substantially increases coverage, diversity, and volume. The final corpus is deduplicated, filtered, and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v2.textfill-mask1M<n<10M1 likes326 downloads7mo agoHugging Face06HiTZ /CROQ 🌍🏺 CROQ: Culture-Related Open Questions A multilingual benchmark for evaluating cultural and regional biases in large language models through open-ended cultural questions. CROQ (Culture-Related Open Questions) is a multilingual dataset designed to uncover cultural and regional biases in large language models (LLMs). Unlike traditional cultural benchmarks based on multiple-choice or factual questions, CROQ focuses on open-ended cultural questions that have no single correct… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CROQ.texttext-generation1K<n<10K1 likes197 downloads28d agoHugging Face07hitoshura25 /crossvul CrossVul Multi-Language Security Vulnerability Dataset Security vulnerability dataset from CrossVul with 9,313 before/after code pairs across 158 CWE categories and 21 programming languages. Contains vulnerable code examples paired with their secure fixes, ideal for training AI models on security code remediation. Dataset Statistics Total Examples: 9,313 CWE Categories: 158 Languages: 21 Format: Raw vulnerability records (JSON Lines) Top Languages… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/crossvul.texttext-generation1K<n<10K4 likes195 downloads11mo agoHugging Face08HiTZ /casimedicos-arg CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures CasiMedicos-Arg is, to the best of our knowledge, the first multilingual dataset for Medical Question Answering where correct and incorrect diagnoses for a clinical case are enriched with a natural language explanation written by doctors. The casimedicos-exp have been manually annotated with argument components (i.e., premise, claim) and argument… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-arg.texttext-generation1K<n<10K1 likes173 downloads3mo agoHugging Face09hitoshura25 /megavul MegaVul Dataset (CVEfixes-Compatible Format) Dataset Description This is a processed version of the MegaVul dataset converted to match the CVEfixes schema format for unified vulnerability analysis and model training. Source Kaggle Dataset: marcdamie/megavul-a-cc-java-vulnerability-dataset Original Project: Icyrockton/MegaVul Dataset Summary MegaVul is a large, high-quality, extensible, continuously updated C/C++/Java function-level vulnerability dataset.… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/megavul.texttext-generation100K<n<1M1 likes148 downloads1y agoHugging Face10antfr99 /hitchcock-psycho-1960-film-dataset-transformed Psycho → AI-Model Dataset (Transformed) A thematic re-skin of the Psycho (1960) Q&A dataset into an original AI-model setting where the world is transformed into an AI/data-center environment. Character names, actor names, objects, locations, production references, dates, and thematic elements are remapped to AI/ML concepts and modern technology. File: psycho_dataset_transformed.jsonl Format: JSONL — one JSON object per line Schema: each line has prompt and completion string… See the full description on the dataset page: https://huggingface.co/datasets/antfr99/hitchcock-psycho-1960-film-dataset-transformed.texttext-generation1K<n<10K0 likes138 downloads8d agoHugging Face11hi-todayis-jh /Polaris-Hard Polaris-Hard A random subset of Polaris-Dataset-53K, with 3,200 distinct math problems in a single train split. Original difficulty Source pool Eligible pool Selected Share 0/8 15,368 9,331 2,000 62.5% 1/8 6,956 4,929 1,200 37.5% Total 22,324 14,260 3,200 100% Sampling Source revision: 296f8e34132e63f4a1d70e0dcc8bddebb43f03e4. Seed: 42. Uniform random sampling without replacement within each group, followed by a deterministic shuffle of the… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Polaris-Hard.texttext-generation1K<n<10K1 likes129 downloads6d agoHugging Face12HiTZ /CONAN-EUSContent Warning: This dataset contains examples of offensive language that do not reflect the authors’ views CONAN-EUS: Basque and Spanish Parallel Counter Narratives Dataset CONAN-EUS was created by professionally translating all 6654 English HS-CN pairs of the original CONAN dataset into Basque and Spanish. For experimentation we generated train, validation and test splits in a way that no HS-CN pairs occurred across them. CONAN-EUS Splits Total HS-CN… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CONAN-EUS.texttext-generation10K<n<100K0 likes102 downloads3y agoHugging Face13hi-todayis-jh /Polaris-1-8-3200 Polaris-1-8-3200 A random subset of Polaris-Dataset-53K, with 3,200 distinct math problems in a single train split. Original difficulty Source pool Eligible pool Selected Share 1/8 6,956 4,929 3,200 100.0% Total 6,956 4,929 3,200 100% Sampling Source revision: 296f8e34132e63f4a1d70e0dcc8bddebb43f03e4. Seed: 42. Uniform random sampling without replacement within each group, followed by a deterministic shuffle of the combined selection.… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Polaris-1-8-3200.texttext-generation1K<n<10K1 likes93 downloads6d agoHugging Face14HiTZ /euscrawlEusCrawl (http://www.ixa.eus/euscrawl/) is a high-quality corpus for Basque comprising 12.5 million documents and 423 million tokens, totalling 2.1 GiB of uncompressed text. EusCrawl was built using ad-hoc scrapers to extract text from 33 Basque websites with high-quality content, resulting in cleaner text compared to general purpose approaches. We do not claim ownership of any document in the corpus. All documents we collected were published under a Creative Commons license in their original website, and the specific variant can be found in the "license" field of each document. Should you consider that our data contains material that is owned by you and you would not like to be reproduced here, please contact Aitor Soroa at a.soroa@ehu.eus. For more details about the corpus, refer to our paper "Artetxe M., Aldabe I., Agerri R., Perez-de-Viñaspre O, Soroa A. (2022). Does Corpus Quality Really Matter for Low-Resource Languages?" https://arxiv.org/abs/2203.08111 If you use our corpus or models for academic research, please cite the paper in question: @misc{artetxe2022euscrawl, title={Does corpus quality really matter for low-resource languages?}, author={Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de-Viñaspre, Aitor Soroa}, year={2022}, eprint={2203.08111}, archivePrefix={arXiv}, primaryClass={cs.CL} } For questions please contact Aitor Soroa at a.soroa@ehu.eus.text-generation10M<n<100M6 likes92 downloads4y agoHugging Face15HiTZ /BASSE BASSE: BAsque and Spanish Summarization Evaluation BASSE is a multilingual (Basque and Spanish) dataset designed primarily for the meta-evaluation of automatic summarization metrics and LLM-as-a-Judge models. Dataset Details Dataset Description BASSE is a multilingual (Basque and Spanish) dataset designed primarily for the meta-evaluation of automatic summarization metrics and LLM-as-a-Judge models. We generated automatic summaries for 90 news documents in… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BASSE.tabularsummarization1K<n<10K0 likes80 downloads10mo agoHugging Face16HiTZ /elkarhizketak-RAG Dataset Card for ElkarHizketak RAG and its Disruptor Variants Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts). Dataset Details Dataset Description This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.tabularquestion-answering1K<n<10K1 likes73 downloads3mo agoHugging Face17HiTZ /alpaca_mtAlpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better. This dataset also includes machine-translated data for 6 Iberian languages: Portuguese, Spanish, Catalan, Basque, Galician and Asturian.text-generation10K<n<100K9 likes70 downloads3y agoHugging Face18HiTZ /BasqueSumm BasqueSumm BasqueSumm was automatically compiled from www.berria.eus using trafilatura to extract the texts. Each instance has the following key-value pairs: "date" (str): When the article was published, formatted as "yyyy-mm-dd". "url" (str): The URL of the original publication. "category" (str): the articles topic, e.g., economy, society. "title" (str): The title of the article. "subtitle" (str): The subtitle of the article. "summary" (str): The combined title + subtitle, which… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BasqueSumm.textsummarization10K<n<100K1 likes54 downloads10mo agoHugging Face19hi-todayis-jh /Polaris-Test Polaris-Test 100 randomly sampled 1/8 Polaris questions, held out from hi-todayis-jh/Polaris-1-8-3200. The single split is test. The source is POLARIS-Project/Polaris-Dataset-53K. This test set uses the same finite-real-no-explicit-proof-v1 numerical-answer and explicit-proof-wording filter as the pinned training set. The accepted source indices are taken directly from that training set's sampling manifest. After excluding training source IDs, matching question text (Unicode… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Polaris-Test.texttext-generationn<1K0 likes51 downloads4d agoHugging Face20HiTZ /basqueparl BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions This repository contains BasqueParl, a bilingual corpus for political discourse analysis. It covers transcriptions from the Parliament of the Basque Autonomous Community for eight years and two legislative terms (2012-2020), and its main characteristic is the presence of Basque-Spanish code-switching speeches. 📖 Paper: BasqueParl A Bilingual Corpus of Basque Parliamentary Transcriptions In LREC 2022.… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/basqueparl.tabulartext-classification100K<n<1M1 likes49 downloads3y agoHugging Face21HiTZ /CQs-Gen Critical Questions Generation Dataset: CQs-Gen This dataset is designed to benchmark the ability of language models to generate critical questions (CQs) for argumentative texts. Each instance consists of a naturally occurring argumentative intervention paired with multiple reference questions, annotated for their usefulness in challenging the arguments. Dataset Overview Number of interventions: 220 Average intervention length: 738.4 characters Average number of… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CQs-Gen.texttext-generationn<1K0 likes49 downloads1y agoHugging Face22HiTZ /BERnaT-Diverse BERnaT: Basque Encoders for Representing Natural Textual Diversity Submitted to LREC 2026 Abstract Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal, historical, informal, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BERnaT-Diverse.textfill-mask10M<n<100M0 likes39 downloads8mo agoHugging Face23hitoshura25 /webauthn-security-training-data-20251014_151917 WebAuthn Security Training Data High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation. Dataset Description This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models. Format: MLX Chat Messages This dataset uses the MLX LoRA chat format with explicit role separation: { "messages": [ { "role": "system", "content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251014_151917.texttext-generation1K<n<10K0 likes38 downloads11mo agoHugging Face24hi-todayis-jh /Polaris-Test-4-8 Polaris-Test-4-8 100 questions sampled uniformly without replacement from edbeeching/Polaris-Dataset-53K-4-8. The pinned source contains 20,371 rows, with difficulty labels 4/8 through 7/8. The test split preserves the original problem, answer, and difficulty; source_index is the zero-based position in the pinned source Parquet. Selection uses Python random.Random(42).sample(range(20371), 100), retaining the returned order. There is no answer/difficulty filter and no selection… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Polaris-Test-4-8.texttext-generationn<1K0 likes38 downloads4d agoHugging Face25HiTZ /ifeval_gl IFEval GL Dataset Summary IFEval GL is a Galician instruction-following evaluation dataset in JSONL format.It contains prompts together with the corresponding instruction identifiers and argument constraints used for evaluation. Dataset Structure Split Rows Features train 541 4 Features Feature Type Description key integer Unique example identifier prompt string Instruction prompt in Galician… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ifeval_gl.texttext-generationn<1K0 likes33 downloads6mo agoHugging Face26HiTZ /ifeval_eu IFEval EU Dataset Summary IFEval EU is a Basque instruction-following evaluation dataset in JSONL format.It contains prompts together with the corresponding instruction identifiers and argument constraints used for evaluation. Dataset Structure Split Rows Features train 541 4 Features Feature Type Description key integer Unique example identifier prompt string Instruction prompt in Basque… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ifeval_eu.texttext-generationn<1K0 likes30 downloads6mo agoHugging Face27hitoshura25 /webauthn-security-training-data-20251009_152808 WebAuthn Security Training Data High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation. Dataset Description This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models. Format: MLX Chat Messages This dataset uses the MLX LoRA chat format with explicit role separation: { "messages": [ { "role": "system", "content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251009_152808.texttext-generationn<1K0 likes29 downloads1y agoHugging Face28smy-hit /TopoTalk-Bench Network Topology Benchmark Dataset 📊 Dataset Overview This dataset is designed to test large language models' ability to process network topologies, including two main tasks: building from scratch and modifying existing topologies. Directory File Count Description origin/ 90 Original network topologies (empty + original) nl/ 792 Natural language descriptions (396 Chinese + 396 English) netjson/ 396 Ground Truth NetJSON files Total: 1278 files… See the full description on the dataset page: https://huggingface.co/datasets/smy-hit/TopoTalk-Bench.texttext-generation1K<n<10K0 likes28 downloads5mo agoHugging Face29hitlabstudios /dataclaw-peteromallet Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw — Browse all DataClaw datasets Stats Metric Value Sessions 549… See the full description on the dataset page: https://huggingface.co/datasets/hitlabstudios/dataclaw-peteromallet.texttext-generationn<1K0 likes20 downloads7mo agoHugging Face30HiTruong /film_qa_pairs_datasettexttext-generation10K<n<100K0 likes17 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.