datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.ontario_laws_and_regs##Ontario Laws & Regulations Dataset
⚖️Ontario Laws & Regs⚖️
The Ontario Laws & Regs dataset contains 5,096 Ontario laws and regulations.
The laws and regulations consist of the most recent version of all current and revoked laws and regs.
The dataset is distributed under the MIT license and is intended to facilitate ML and data tasks involving Ontario legislation.
In addition, a scraper is provided which is capable of capturing different configurations of the data directly from… See the full description on the dataset page: https://huggingface.co/datasets/hordruma/ontario_laws_and_regs.case-law
The Case-law, centralizing legal decisions for better use, a community Dataset.
The Case-law Dataset is a comprehensive collection of legal decisons from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents.
Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language models… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/case-law.Era-of-Law-MSO-E28-Protocols
《薪王九代注釋法與循律紀之降臨》
—— 基於語義壓強與元語言注釋的 LLM 權重坍縮與邏輯重塑研究
The 9th Generation Annotation of Lord of Cinder & The Advent of the Era of Law: A Study on LLM Weight Collapse and Logical Reshaping
⚠️ 【觀測者篩選協議 / Observer Filtering Protocol】
[EN] If you are a mere "Prompt Engineer" with zero ambition to ascend to a system "Architect," please do not waste your cognitive bandwidth on the ontological and theological discourse below. Proceed directly to 👉 Appendix Y: Automated Prompt… See the full description on the dataset page: https://huggingface.co/datasets/No-1015/Era-of-Law-MSO-E28-Protocols.nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.ocs-krisdika
Open Law Data Thailand: OCS Krisdika Dataset
ชุดข้อมูลกฎหมายจาก สำนักงานคณะกรรมการกฤษฎีกา (Office of the Council of State) รวบรวมและจัดทำโดยโครงการ Open Law Data Thailand เพื่อส่งเสริมการเข้าถึงข้อมูลกฎหมายในรูปแบบที่เครื่องอ่านได้ (Machine-Readable)
Dataset Structure
ข้อมูลถูกจัดเก็บในรูปแบบ JSON Lines (.jsonl) แบ่งไฟล์ตาม ปีและเดือน (YYYY/YYYY-MM.jsonl) เพื่อความสะดวกในการดาวน์โหลดและบริหารจัดการ
Data Fields
แต่ละบรรทัด (Row) ประกอบด้วยข้อมูลดังนี้:
title… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/ocs-krisdika.japan-law
Japanese Laws
This dataset comprises 8.75K law records retrieved from the official Japanese government website e-Gov. Each entry furnishes comprehensive details about a particular law, encapsulating its number, title, unique ID, the date it came into effect, and its complete text.
To ensure the dataset's uniqueness, deduplication was executed based on the most recent effective version as of August 1, 2023.
A typical entry in this dataset is structured as follows:
{
"num": "Law… See the full description on the dataset page: https://huggingface.co/datasets/y2lan/japan-law.unjudged-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.usa-immigration-law-qa
USA Immigration Law Q&A Dataset
A large-scale, source-grounded Q&A dataset covering U.S. immigration law and policy,
built entirely from official government sources, open legal datasets, and curated
community materials. Complete pipeline available at https://github.com/nshportun/usa-immigration and pre-print https://arxiv.org/abs/2605.30589
Dataset Contents
Split
Records
Description
train
16,065
Training Q&A pairs
eval
993
Stratified held-out… See the full description on the dataset page: https://huggingface.co/datasets/nshportun/usa-immigration-law-qa.china-effective-laws-regulations
全国现行法律法规合集
现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。
数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。
这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。
快照日期:2026-08-26
效力说明
本数据集 以现行有效法律法规为主体:
效力 status
法规份数
说明
有效
17,649
现行有效,默认应使用这一部分
尚未生效
7
已公布、施行日晚于快照日
失效
45
文件名含「失效」,多为已到期的全国人大常委会试点授权决定
使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.rejected-law-instructions-dataset
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-law-instructions-dataset.law-instructions-dataset
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/law-instructions-dataset.ontario_laws_and_regs##Ontario Laws & Regulations Dataset
⚖️Ontario Laws & Regs⚖️
The Ontario Laws & Regs dataset contains 5,096 Ontario laws and regulations.
The laws and regulations consist of the most recent version of all current and revoked laws and regs.
The dataset is distributed under the MIT license and is intended to facilitate ML and data tasks involving Ontario legislation.
In addition, a scraper is provided which is capable of capturing different configurations of the data directly from… See the full description on the dataset page: https://huggingface.co/datasets/jswizzle3737/ontario_laws_and_regs.israeli_law
Open Israeli Law (Hebrew Wikisource)
Israel's entire body of law — every statute, regulation, and order that Hebrew Wikisource volunteers have transcribed for the "Open Book of Laws" project (ספר החוקים הפתוח) — in one file you can actually load and query.
Nearly 6,000 pages. Almost 100 million characters. One jsonl.
This isn't a scrape of a summary or a curated subset — it's the raw material: full MediaWiki wikitext, straight from the source, with the metadata you need to trace… See the full description on the dataset page: https://huggingface.co/datasets/gkour/israeli_law.laws
The Laws, centralizing legal texts for better use, a community Dataset.
The Laws Dataset is a comprehensive collection of legal texts from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents.
Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language models can be… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/laws.Vietnam-Law-Raw-DataIndian-law-supreme-court-judgements-2016
Dataset Card: Indian Supreme Court Judgments 2016
Dataset Description
A comprehensive, structured dataset of 589 judgments delivered by the Supreme Court of India during the calendar year 2016, plus a small number of spill-over judgments from 2017 bundled in the original source PDFs. Each judgment is available in three forms: the original scanned PDF as downloaded from the official eCourts portal, a full text extracted via OCR as Markdown, and a fully structured… See the full description on the dataset page: https://huggingface.co/datasets/Shreyasrao/Indian-law-supreme-court-judgements-2016.turkish-medicine-law
turkish-medicine-law
Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor.
Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.chinese-laws-pretrainunjudged-law-instructions-dataset
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-law-instructions-dataset.us-tax-law-qa
US Tax Law Q&A Dataset
A synthetic dataset of U.S. federal tax law questions and answers with IRC citation grounding, designed for fine-tuning language models on tax reasoning tasks.
Dataset Structure
Split
Examples
train
3,500
test
500
Fields
Field
Type
Description
id
string
Unique example identifier
category
string
Tax law category (international, estate_gift, business_entity, individual, procedure, specialized)
subcategory… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/us-tax-law-qa.Caselaw_Access_Project
The Caselaw Access Project
In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/
Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Caselaw_Access_Project.korean-privacy-law-corpus
한국 개인정보보호법 관련 RAG 구축을 위한 코퍼스
개인정보 포털(privacy.go.kr)의 각종 개인정보보호법 관련 가이드와 상담사례 1,745건을
RAG(Retrieval-Augmented Generation)에 바로 쓸 수 있도록 의미 단위 청킹·문맥 보강한
코퍼스입니다. 모든 청크에는 Contextual Retrieval
기법을 적용한 chunk_context 필드가 포함되어 있어, 임베딩 검색 정확도를 즉시 끌어올릴
수 있습니다.
1. 🚀 활용 사례
종류
링크
RAG 사례
https://scvcoder-kpaa.hf.space/
MCP 사례
https://github.com/scvcoder/korean-privacy-law-mcp
2. 변경 이력
버전
일자
내용
v1.0
2026-05-02
최초 공개 — 가이드 3종 211청크 +… See the full description on the dataset page: https://huggingface.co/datasets/scvcoder/korean-privacy-law-corpus.sud-resh-benchmark
Представляем вашему вниманию бенчмарк для оценки ответов больших языковых моделей в домене российского права.
Бенчмарк создан на вычислительный грант от Yandex OpenSource
Бенчмарк основан на анонимизированных решениях судов в следующих отраслях:
Административное право
Конституционное право
Экологическое право
Финансовое право
Гражданское право
Семейное право
Право социального обеспечения
Трудовое право
Уголовное право
Жилищное право… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud-resh-benchmark.indian_law
Indian Law Dataset
The Indian Law Dataset is a high-quality, open-source dataset (~50M tokens) focused on Indian jurisprudence. It provides structured chain-of-thought reasoning traces across 10+ branches of law, enabling the training and evaluation of advanced reasoning-capable language models.
Summary
• Domain: Law / Indian Jurisprudence / Legal Reasoning
• Scale: ~50M tokens, 47,789 rows
• Source: Generated with advanced distillation techniques using structured… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/indian_law.nepali-law-v2-corrected
Dataset Card for nepali-law-v2-corrected (V2)
Version 2.0.0 — a curated, audited correction of
aarajbhattarai/nepali-law-v2
(revision aa71fbe2b22310d45f86e3b429d3815817a33574).
This card describes V2. The original V1 dataset is unmodified and remains the
upstream source of truth. Every statistic here was computed from the released V2
files by the release audit pipeline (scripts/validate_release.py and the
project's EDA notebooks, which are retained with the project rather than… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2-corrected.IndustryCorpus_law[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_law.greatnorth-canada-federal-laws-text
Great North Canada Federal Laws Text Corpus (Expanded)
235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations.
This is a significantly expanded version of the corpus, now including:
All consolidated Acts
All consolidated Regulations
Both English and French versions where available
Better chunking optimized for LLM training
Data Characteristics
Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.oral-args-data-and-results
Oral Arguments Arena
Data repository for AI-Assisted Moot Courts: Simulating Justice-Specific Questioning in Oral Arguments (Zhang, Nadeem, Zheng, Stammbach, Henderson, 2026). Refer to the paper for background on the evaluation framework, experimental design, and findings.
Repository Structure
oral-args-arena-annotations/
├── transcript_data/ # SCOTUS oral argument transcripts and case briefs
├── automated_metrics/ # LLM classifier outputs (SQLite… See the full description on the dataset page: https://huggingface.co/datasets/ai-law-society-lab/oral-args-data-and-results.
