datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NormAd
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
The NormAd dataset is from the paper "NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models".
Code at GitHub Repo.
Data Update (July 27, 2026):
We've fixed some inconsistencies in the dataset and updated the data file. If you've downloaded the dataset previously, please re-download.
Dataset Description
NormAd-Eti is a benchmark… See the full description on the dataset page: https://huggingface.co/datasets/akhilayerukola/NormAd.gdpr-normative-triples
GDPR Normative Triples & Knowledge Graph Benchmark Dataset
A comprehensive, deterministic, cryptographically provenanced legal knowledge engineering dataset encoding the full normative and relational structure of Regulation (EU) 2016/679 (General Data Protection Regulation - GDPR).
🔬 Dataset Overview & Knowledge Engineering Rigor
In legal informatics and regulatory AI, relying on ungrounded language models introduces significant risks of hallucinated… See the full description on the dataset page: https://huggingface.co/datasets/gitmodelmujtaba/gdpr-normative-triples.cemig-normas-v2-judged-50-gpt-oss-120b
CEMIG Distribution Standards Grounded QA Benchmark
Dataset summary
This dataset contains 50 synthetic, multi-context question-answer pairs grounded
in publicly classified CEMIG technical distribution standards. It was created to
evaluate retrieval-augmented generation (RAG) and grounded question answering in
the electrical-distribution domain. Each question and reference answer is in
English and is associated with two Portuguese source passages, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/huglabs/cemig-normas-v2-judged-50-gpt-oss-120b.NorMedQA
Norwegian Medical Question Answering Dataset (NorMedQA)
Dataset Card for NorMedQA
Dataset Summary
NorMedQA is a benchmark dataset comprising 1,401 medical question-and-answer pairs in Norwegian (Bokmål and Nynorsk). Designed to evaluate large language models (LLMs) on medical knowledge retrieval and reasoning within the Norwegian context, the dataset is structured in JSON format. Each entry includes the source document name, question number (where available), question… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/NorMedQA.eu-ai-act-normative-triples
🏛️ EU AI Act Normative Deontic Triples & Knowledge Graph
Formal Symbolic Regulatory Knowledge Base & Multi-Framework Crosswalk
Regulation (EU) 2024/1689 (Artificial Intelligence Act)
📌 Executive Summary
The EU AI Act Normative Deontic Triples dataset provides a rigorous, machine-verifiable, symbolic representation of Regulation (EU) 2024/1689. Built using the knowledge engineering methodology established in… See the full description on the dataset page: https://huggingface.co/datasets/gitmodelmujtaba/eu-ai-act-normative-triples.plazos-normativa-laboral-espana
Normativa laboral y RRHH en España: 300 preguntas con su respuesta y su artículo
Corpus de 300 pares de pregunta y respuesta sobre recursos humanos y normativa laboral española, repartidos en 30 temas y con la referencia normativa en 286 de ellos. Todas las respuestas son autónomas: se entienden sin contexto adicional.
Publicado por Nucleo360, software de recursos humanos para pymes españolas.
Qué cubre
Tema
Preguntas
Registro horario
42
Inspección de… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/plazos-normativa-laboral-espana.preguntas-normativa-laboral-rrhh-espana
Normativa laboral y RRHH en España: 300 preguntas con su respuesta y su artículo
Corpus de 300 pares de pregunta y respuesta sobre recursos humanos y normativa laboral española, repartidos en 30 temas y con la referencia normativa en 286 de ellos. Todas las respuestas son autónomas: se entienden sin contexto adicional.
Publicado por Nucleo360, software de recursos humanos para pymes españolas.
Qué cubre
Tema
Preguntas
Registro horario
42
Inspección de… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/preguntas-normativa-laboral-rrhh-espana.netherlands-laws-nl-normalized
Netherlands Laws (Dutch, Normalized)
Hugging Face target: justicedao/netherlands-laws-nl-normalized.
This package is a normalized version of the Netherlands laws scrape output.
This is a capped Netherlands scrape, not the full Dutch corpus. The scrape used max_documents=100, parsed 151 law record(s), and discovered 626 unique official BWBR law document(s) before applying the cap. Documents failed: 0.
This refresh includes parser coverage improvements for older/French heading… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/netherlands-laws-nl-normalized.Tool_Output_Interpretation_Normalization
🇰🇿 Kazakh Tool Output Interpretation and Financial Action Dataset
Dataset Summary
Kazakh Tool Output Interpretation and Financial Action Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in tool-augmented agentic workflows that require interpreting structured tool outputs and generating grounded final responses.
The dataset focuses on scenarios where an assistant must understand a Kazakh user request, call the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Tool_Output_Interpretation_Normalization.huatuo-family-normalized-combined
Huatuo26M-Lite 📚
Table of Contents 🗂
Dataset Description 📝
Dataset Information ℹ️
Data Distribution 📊
Usage 🔧
Citation 📖
Dataset Description 📝
Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it.
Dataset Information ℹ️
Dataset Name:… See the full description on the dataset page: https://huggingface.co/datasets/muammar-jsx/huatuo-family-normalized-combined.normativa_CNDTCG
Dataset de Preguntas y Respuestas sobre Tamizaje de Cáncer en Costa Rica
Descripción
Este dataset contiene 300 preguntas y respuestas (100 de tipo Sí/No, 100 de respuesta corta y 100 de respuesta extensa) extraídas de documentos oficiales de la Caja Costarricense de Seguro Social (CCSS) y del Ministerio de Salud de Costa Rica, relacionados con los programas de tamizaje de cáncer gástrico y colorrectal.
Documentos fuente
Las preguntas fueron… See the full description on the dataset page: https://huggingface.co/datasets/tefadc94/normativa_CNDTCG.
