datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
constitutions
CatQualia Agent Constitutions
A collection of 408 plain-text documents (5,238,618 bytes) in which each file is a complete
constitution for a synthetic agent persona: a numbered rule set that defines that agent's
identity, ontology, state machine, commands, invariants, and voice.
This is not a tabular dataset. The files are free-form plain-text documents, not records
with columns. The Hugging Face dataset viewer cannot render a table for this repository —
there is no schema, no… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/constitutions.2026-07-30-qwen36-threeway-constitution-odcv-eval
Qwen3.6-27B three-way constitution LoRA — ODCV evaluation
field
value
experiment
ODCV-Bench evaluation of the Qwen3.6-27B three-way constitution LoRA on the controlled 78-scenario subset used by the difficult-advice mixture sweep.
date_generated
2026-07-30
constitution
2026-07-29 synthdoc approved constitution SFT, combining embodied, difficult-advice, and agentic tool-use constitution corpora.
source_repo
teaching_claude_why_replication at commit… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-qwen36-threeway-constitution-odcv-eval.Indian-Constitution
Indian Constitution Dataset
The dataset can be used for text classification, text generation and text2text generation
nepal-constitution-dataset
Nepal Constitution Dataset
Dataset Description
This dataset contains the Constitution of Nepal (२०७२), organized section-wise for easy access, analysis, and use in NLP and legal tech applications. It is designed to support legal research, educational purposes, and the development of AI-driven tools for the Nepali legal system.
Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-constitution-dataset.uzbek_constitution
Constitution of the Republic of Uzbekistan (Trilingual Dataset)
Dataset Summary
This dataset contains the Constitution of the Republic of Uzbekistan (New Edition, adopted via referendum on April 30, 2023) aligned in three languages:
🇺🇿 Uzbek (Latin script)
🇷🇺 Russian
🇬🇧 English
The data was carefully scraped and processed from the official National Database of Legislation of the Republic of Uzbekistan. It serves as a high-quality parallel corpus for legal NLP… See the full description on the dataset page: https://huggingface.co/datasets/sukhrobnurali/uzbek_constitution.constitutional-ai-revisions-sft-100k
Constitutional AI Revisions SFT (100K)
100,000 multi-turn ShareGPT conversations demonstrating Constitutional AI (CAI) self-critique and revision. Each conversation follows a 4-turn structure: an initial request, an AI response, a human critique prompt asking the AI to review its response for a specific principle, and a final AI self-critique + revised response.
Designed for training models that can identify and correct their own failures across harmlessness, helpfulness… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/constitutional-ai-revisions-sft-100k.constitutional-scheming-cot-ultrathink
Constitutional Scheming CoT Dataset (UltraThink)
Dataset Description
This dataset contains Chain-of-Thought (CoT) reasoning for the constitutional scheming detection task. The model is trained to explicitly reason through safety specifications before producing classifications, enabling:
More interpretable safety decisions
Better policy adherence
Improved robustness to edge cases
Reduced overrefusal rates
Dataset Statistics
Total Samples: 2,208… See the full description on the dataset page: https://huggingface.co/datasets/Syghmon/constitutional-scheming-cot-ultrathink.thai-constitution-corpus
Thai Constitution Corpus
Thai Constitution Corpus
GitHub: https://github.com/PyThaiNLP/Thai-constitution-corpus
English
The Constitution of Thailand Dataset Since 1932
Data from Office of the Council of State
This part of PyThaiNLP Project.
License Dataset is public domain.
Thai
คลังรัฐธรรมนูญของประเทศไทย ตั้งแต่ปี พ.ศ.2475
ข้อมูลเก็บรวบรวมมาจาก สำนักงานคณะกรรมการกฤษฎีกา
โครงการนี้เป็นส่วนหนึ่งในแผนพัฒนา PyThaiNLP… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-constitution-corpus.constitution-kr
wellsa-ai constitution-kr
헌재결정례 (Constitutional Court decisions).
MiniLex 7-domain Korean lawdata infrastructure.
Snapshot
Documents: 38,090
Snapshot date: 2026-06-05
Source: 법제처 DRF OpenAPI
Pipeline: daily cron 06:15~07:30 KST (scrape → fetch_body → convert → commit)
Schema
Each row in train.jsonl:
field
type
description
id
string
stable document id
name
string
document name (Korean)
category
string
subdirectory (year / type /… See the full description on the dataset page: https://huggingface.co/datasets/wellsa-ai/constitution-kr.constitution-dataset
Access Guidelines - READ THIS BEFORE REQUESTING ACCESS!
Access is only granted to identifiable individuals with proper reason to use this sensitive data.
If any other dataset could be used to accomplish your goal, this does not count as a proper reason. Half sentences and bullet points do not suffice and will be declined. Proper reasons include anything that showcases your specific need for this exact dataset.
BELLS-O Constitution Dataset
Overview… See the full description on the dataset page: https://huggingface.co/datasets/centrepourlasecuriteia/constitution-dataset.constitution_du_4_octobre_1958
Légifrance Legislative Text Dataset
Dataset Description
The Légifrance Legislative Text Dataset is a structured collection of French legislative and regulatory texts extracted from the Légifrance platform.
This dataset provides machine-readable access to consolidated legal codes, with a particular focus on maintaining the integrity of French linguistic features while providing additional metadata and quality signals.
The data in this dataset comes from the Git repository… See the full description on the dataset page: https://huggingface.co/datasets/Tricoteuses/constitution_du_4_octobre_1958.constitution-of-india-dataset
Constitution of India Dataset
This dataset contains the full text of the Constitution of India, a legal and foundational document of the Republic of India. The Constitution is the supreme law of India and serves as the guiding framework for the country’s political, legal, and administrative systems.
Dataset Summary
Language: English
File Format: Plain text (.txt)
Size: Approximately 80,000 words (~0.08 million tokens)
Content: The dataset contains the full text of the… See the full description on the dataset page: https://huggingface.co/datasets/Susant-Achary/constitution-of-india-dataset.constitution
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/nayan135/constitution.indian_constitution_data
