CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HiTZ /casimedicos-exp Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments for the correct answer but also arguments to explain why the remaining possible answers are incorrect. This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation. The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.tabulartext-generation1K<n<10K4 likes1.9k downloads3y agoHugging Face02HiTZ /MedExpQA MexExpQA: Multilingual Benchmarking of Medical QA with reference gold explanations and Retrieval Augmented Generation (RAG) We present a new multilingual parallel medical benchmark, MedExpQA, for the evaluation of LLMs on Medical Question Answering. This benchmark can be used for various NLP tasks including: Medical Question Answering or Explanation Generation. Although the design of MedExpQA is independent of any specific dataset, for the first version of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MedExpQA.tabulartext-generation1K<n<10K9 likes1.9k downloads2y agoHugging Face03HiTZ /latxa-corpus-v1.1 Latxa Corpus v1.1 This is the training corpus of Latxa v1.1, a family of large language models for Basque based on Llama 2. 💻 Repository: https://github.com/hitz-zentroa/latxa 📒 Blog Post: Latxa: An Open Language Model and Evaluation Suite for Basque 📖 Paper: Latxa: An Open Language Model and Evaluation Suite for Basque 📧 Point of Contact: hitz@ehu.eus 📌 Notice As of February 13th 2026, this repository reflects a curated version of the original dataset. Some data… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v1.1.textfill-mask1M<n<10M2 likes568 downloads7mo agoHugging Face04HiTZ /latxa-corpus-v2 Latxa Corpus v2 📧 Point of Contact: hitz@ehu.eus Dataset Summary Curated by: HiTZ Research Center & IXA Research group (University of the Basque Country UPV/EHU) Language(s): eu-ES Latxa Corpus v2 is a large-scale monolingual Basque corpus, created by combining curated crawls, public datasets, institutional data, and newly collected resources. Compared to v1.1, it substantially increases coverage, diversity, and volume. The final corpus is deduplicated, filtered, and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v2.textfill-mask1M<n<10M1 likes320 downloads7mo agoHugging Face05antfr99 /hitchcock-psycho-1960-film-dataset-transformed Psycho → AI-Model Dataset (Transformed) A thematic re-skin of the Psycho (1960) Q&A dataset into an original AI-model setting where the world is transformed into an AI/data-center environment. Character names, actor names, objects, locations, production references, dates, and thematic elements are remapped to AI/ML concepts and modern technology. File: psycho_dataset_transformed.jsonl Format: JSONL — one JSON object per line Schema: each line has prompt and completion string… See the full description on the dataset page: https://huggingface.co/datasets/antfr99/hitchcock-psycho-1960-film-dataset-transformed.texttext-generation1K<n<10K0 likes140 downloads9d agoHugging Face06HiTZ /elkarhizketak-RAG Dataset Card for ElkarHizketak RAG and its Disruptor Variants Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts). Dataset Details Dataset Description This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.tabularquestion-answering1K<n<10K1 likes75 downloads3mo agoHugging Face07HiTZ /BasqueSumm BasqueSumm BasqueSumm was automatically compiled from www.berria.eus using trafilatura to extract the texts. Each instance has the following key-value pairs: "date" (str): When the article was published, formatted as "yyyy-mm-dd". "url" (str): The URL of the original publication. "category" (str): the articles topic, e.g., economy, society. "title" (str): The title of the article. "subtitle" (str): The subtitle of the article. "summary" (str): The combined title + subtitle… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BasqueSumm.textsummarization10K<n<100K1 likes55 downloads7h agoHugging Face08HiTZ /CQs-Gen Critical Questions Generation Dataset: CQs-Gen This dataset is designed to benchmark the ability of language models to generate critical questions (CQs) for argumentative texts. Each instance consists of a naturally occurring argumentative intervention paired with multiple reference questions, annotated for their usefulness in challenging the arguments. Dataset Overview Number of interventions: 220 Average intervention length: 738.4 characters Average number of… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CQs-Gen.texttext-generationn<1K0 likes49 downloads1y agoHugging Face09HiTZ /BERnaT-Diverse BERnaT: Basque Encoders for Representing Natural Textual Diversity Submitted to LREC 2026 Abstract Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal, historical, informal, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BERnaT-Diverse.textfill-mask10M<n<100M0 likes37 downloads8mo agoHugging Face10HiTZ /ifeval_gl IFEval GL Dataset Summary IFEval GL is a Galician instruction-following evaluation dataset in JSONL format.It contains prompts together with the corresponding instruction identifiers and argument constraints used for evaluation. Dataset Structure Split Rows Features train 541 4 Features Feature Type Description key integer Unique example identifier prompt string Instruction prompt in Galician… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ifeval_gl.texttext-generationn<1K0 likes31 downloads6mo agoHugging Face11HiTZ /ifeval_eu IFEval EU Dataset Summary IFEval EU is a Basque instruction-following evaluation dataset in JSONL format.It contains prompts together with the corresponding instruction identifiers and argument constraints used for evaluation. Dataset Structure Split Rows Features train 541 4 Features Feature Type Description key integer Unique example identifier prompt string Instruction prompt in Basque… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ifeval_eu.texttext-generationn<1K0 likes28 downloads6mo agoHugging Face12hitoshura25 /webauthn-security-training-data-20251014_151917 WebAuthn Security Training Data High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation. Dataset Description This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models. Format: MLX Chat Messages This dataset uses the MLX LoRA chat format with explicit role separation: { "messages": [ { "role": "system", "content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251014_151917.texttext-generation1K<n<10K0 likes27 downloads1y agoHugging Face13hitoshura25 /webauthn-security-training-data-20251009_152808 WebAuthn Security Training Data High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation. Dataset Description This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models. Format: MLX Chat Messages This dataset uses the MLX LoRA chat format with explicit role separation: { "messages": [ { "role": "system", "content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251009_152808.texttext-generationn<1K0 likes24 downloads1y agoHugging Face14hitlabstudios /dataclaw-peteromallet Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw — Browse all DataClaw datasets Stats Metric Value Sessions 549… See the full description on the dataset page: https://huggingface.co/datasets/hitlabstudios/dataclaw-peteromallet.texttext-generationn<1K0 likes20 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.