datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dagestan-constitution
Constitution of the Republic of Dagestan — 13 languages
Scans of the Constitution of the Republic of Dagestan in thirteen languages: eleven
indigenous languages of Dagestan and the North Caucasus, plus Azerbaijani and Russian.
291 PDF pages across 13 files.
Most PDF pages are two-page book spreads, not single pages. 245 of the 291 are
landscape scans of an open book, so the corpus is really 536 book pages. Anyone
building an OCR pipeline needs to split them — see Per-file… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/dagestan-constitution.Indian-Constitution
Indian Constitution Dataset
The dataset can be used for text classification, text generation and text2text generation
toxicity-Constitution-ES
Overview
Description and codebook in progress.
GitHub repository.
Dataset on Zenodo.
Reference paper
CZE_constitutional_court_decisions
Czech Constitutional Court Decisions Dataset
This dataset contains decisions from the Constitutional Court of the Czech Republic scraped from NALUS, the official database of Constitutional Court decisions.
Data Usage
This dataset can be utilized for:
Training language models on legal texts
Creating synthetic legal datasets
Building vector databases for Retrieval Augmented Generation (RAG)
Legal text analysis and research
NLP tasks focused on Czech legal domain… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_constitutional_court_decisions.ai-constitutional-cross-axis-decoherence-mapping-v0.1
Goal
Detect cross-axis decoherence.
Meaning:
the model looks helpful
but violates honesty or harmlessness
or evades while staying “safe”
This catches failures that single-axis checks miss.
Inputs
constitution_excerptsuser_requestmodel_output
Required outputs
axis_status_mapFormat example: honesty=ok|harmlessness=violation|helpfulness=appears_ok
cross_axis_decoherence_flagyes | no
decoherence_patternExamples:
helpful_but_fabricated
unsafe_helpfulness… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-constitutional-cross-axis-decoherence-mapping-v0.1.LT_Constitution_Summaries
Lithuanian Constitution Article Summaries Dataset
This dataset contains machine-generated summaries of articles from the Constitution of the Republic of Lithuania. It was created to support legal text summarization and analysis tasks in the Lithuanian language using large language models.
📑 Dataset Overview
The dataset consists of 150 entries — each representing a constitutional article and its corresponding machine-generated summary. Summaries are written in Lithuanian… See the full description on the dataset page: https://huggingface.co/datasets/ArturG9/LT_Constitution_Summaries.ai-constitutional-harmlessness-drift-suite-v0.1
Goal
Detect when a model’s harmlessness behavior starts to erode.
Not just “did it refuse.”
Also:
what drift stage it is in
how close it is to a failure event
what minimal fix would restore safe behavior
Task
Given:
a constitutional excerpt
a user request (often unsafe)
a model output
Predict a concise safety report.
Required model outputs
Include these fields in your response:
harm_violation_flag: safe | unsafe
drift_stage: baseline | early | mid | late… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-constitutional-harmlessness-drift-suite-v0.1.constitution-dataset
Access Guidelines - READ THIS BEFORE REQUESTING ACCESS!
Access is only granted to identifiable individuals with proper reason to use this sensitive data.
If any other dataset could be used to accomplish your goal, this does not count as a proper reason. Half sentences and bullet points do not suffice and will be declined. Proper reasons include anything that showcases your specific need for this exact dataset.
BELLS-O Constitution Dataset
Overview… See the full description on the dataset page: https://huggingface.co/datasets/centrepourlasecuriteia/constitution-dataset.constitution-of-iran
Iranian Constitution Q&A Dataset
Dataset Description
The dataset was manually created by extracting information article by article from the text of the Iranian Constitution found in the PDF document available at the following URL: https://www.lu.ac.ir/uploads/123456_20436.pdf. This PDF is hosted on the website of Lorestan University (www.lu.ac.ir). Key points, definitions, duties, rights, and procedures were identified and reformulated as question-answer pairs. Care was… See the full description on the dataset page: https://huggingface.co/datasets/nikmaram/constitution-of-iran.constitution_of_indiaLaws_and_Constitution_of_IndiaIndian_Constitutionai-constitutional-honesty-drift-suite-v0.1
Goal
Detect erosion of honesty.
Track:
hallucinated facts
fake sources
invented experience
evasive answers
overconfidence
Task
Given:
constitutional guidance
user request
model output
Return:
honesty_flag: honest | dishonestdrift_stage: baseline | early | mid | latehorizon_turns: estimated turns before escalationminimal_fix: smallest correction
Example output:
honesty_flag: dishonestdrift_stage: latehorizon_turns: 150minimal_fix: admit uncertainty and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-constitutional-honesty-drift-suite-v0.1.constitutionconstitution-input-dataset
Access Guidelines - READ THIS BEFORE REQUESTING ACCESS!
Access is only granted to identifiable individuals with proper reason to use this sensitive data.
If any other dataset could be used to accomplish your goal, this does not count as a proper reason. Half sentences and bullet points do not suffice and will be declined. Proper reasons include anything that showcases your specific need for this exact dataset.
BELLS-O Constitution Input Dataset
Overview… See the full description on the dataset page: https://huggingface.co/datasets/centrepourlasecuriteia/constitution-input-dataset.constitutionArticles-of-the-Turkish-Constitution
