datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.Indian-Laws
Dataset Card for Indian Laws
This is a comprehensive collection of primary legal documents pertinent to the Indian legal system.
It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law.
open-india-law
Open India Law
Open, structured Indian primary law - plus the scrapers that build it.
Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15
tribunals and regulators, and Central, State and Union Territory legislation down to the
individual section. Normalized to one schema, exclusively from official government sources.
Volume
Period
Court judgments
12,848,644
1950 to 2025
Tribunal and regulator matters
813,168
1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.pile-of-lawWe curate a large corpus of legal and administrative data. The utility of this data is twofold: (1) to aggregate legal and administrative data sources that demonstrate different norms and legal standards for data filtering; (2) to collect a dataset that can be used in the future for pretraining legal-domain language models, a key direction in access-to-justice initiatives.ontario_laws_and_regs##Ontario Laws & Regulations Dataset
⚖️Ontario Laws & Regs⚖️
The Ontario Laws & Regs dataset contains 5,096 Ontario laws and regulations.
The laws and regulations consist of the most recent version of all current and revoked laws and regs.
The dataset is distributed under the MIT license and is intended to facilitate ML and data tasks involving Ontario legislation.
In addition, a scraper is provided which is capable of capturing different configurations of the data directly from… See the full description on the dataset page: https://huggingface.co/datasets/hordruma/ontario_laws_and_regs.Indian-Laws
Dataset Card for Indian Laws
This is a comprehensive collection of primary legal documents pertinent to the Indian legal system.
It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law.
ipfs_dominicanrepublic_laws
Laws of Dominican Republic
Research snapshot of official legislation collected from Consultoria Juridica / consultoria.gov.do.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
Consultoria Juridica / consultoria.gov.do
Collector
scrapers/collect_do.py
Laws / instruments
5669
Articles
278085
Language
es
Jurisdiction… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_dominicanrepublic_laws.ipfs_angola_laws
Laws of Angola
Research snapshot of official legislation collected from Diario da Republica / GUE / Portal do Governo.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
Diario da Republica / GUE / Portal do Governo
Collector
scrapers/collect_ao.py
Laws / instruments
5480
Articles
53019
Language
pt
Jurisdiction
Angola… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_angola_laws.ipfs_ethiopia_laws
Laws of Ethiopia
Research snapshot of official legislation collected from Federal Negarit Gazette.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
Federal Negarit Gazette
Collector
scrapers/collect_et.py
Laws / instruments
5164
Articles
19221
Language
am
Jurisdiction
Ethiopia
License
et-negarit… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_ethiopia_laws.ipfs_albania_laws
Laws of Albania
Research snapshot of official legislation collected from QPZ/QBZ Fletorja Zyrtare PDFs.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
QPZ/QBZ Fletorja Zyrtare PDFs
Collector
scrapers/collect_al.py
Laws / instruments
5870
Articles
108624
Language
sq
Jurisdiction
Albania
License
al-qbz-fletore… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_albania_laws.ipfs_azerbaijan_laws
Laws of Azerbaijan
Research snapshot of official legislation collected from e-qanun.az downloadDetailPdf.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
e-qanun.az downloadDetailPdf
Collector
scrapers/collect_az.py
Laws / instruments
7101
Articles
6241
Language
az
Jurisdiction
Azerbaijan
License
az-eqanun… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_azerbaijan_laws.ipfs_iran_laws
Laws of Iran
Research snapshot of official legislation collected from DOTIC dotic.ir.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
DOTIC dotic.ir
Collector
scrapers/collect_ir.py
Laws / instruments
5904
Articles
11374
Language
fa
Jurisdiction
Iran
License
ir-dotic
Contents
data/laws.parquet —… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_iran_laws.ipfs_bolivia_laws
Laws of Bolivia
Research snapshot of official legislation collected from Gaceta Oficial de Bolivia.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
Gaceta Oficial de Bolivia
Collector
scrapers/collect_bo.py
Laws / instruments
5335
Articles
39159
Language
es
Jurisdiction
Bolivia
License
bo-gaceta-oficial… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_bolivia_laws.ipfs_tunisia_laws
Laws of Tunisia
Research snapshot of official legislation collected from Journal Officiel (IORT/JORT; iort.gov.tn / ocr.jort.tn).
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
Journal Officiel (IORT/JORT; iort.gov.tn / ocr.jort.tn)
Collector
scrapers/collect_tn.py
Laws / instruments
6395
Articles
84161
Language
fr… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_tunisia_laws.ai-lawyeripfs_serbia_laws
Laws of Serbia
Research snapshot of official legislation collected from PIS + parlament.gov.rs zakoni PDFs.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
PIS + parlament.gov.rs zakoni PDFs
Collector
scrapers/collect_rs.py
Laws / instruments
6348
Articles
60111
Language
sr
Jurisdiction
Serbia
License… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_serbia_laws.ipfs_zambia_laws
Laws of Zambia
Research snapshot of official legislation collected from Parliament of Zambia / ZambiaLII.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
Parliament of Zambia / ZambiaLII
Collector
scrapers/collect_zm.py
Laws / instruments
5004
Articles
10228
Language
en
Jurisdiction
Zambia
License
zm-parliament… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_zambia_laws.ipfs_malawi_laws
Laws of Malawi
Research snapshot of official legislation collected from MalawiLII / Parliament of Malawi.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
MalawiLII / Parliament of Malawi
Collector
scrapers/collect_mw.py
Laws / instruments
5008
Articles
4426
Language
en
Jurisdiction
Malawi
License
mw-malawilii… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_malawi_laws.ipfs_laos_laws
Laws of Laos
Research snapshot of official legislation collected from Lao Official Gazette.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
Lao Official Gazette
Collector
scrapers/collect_la.py
Laws / instruments
5111
Articles
34506
Language
lo
Jurisdiction
Laos
License
la-official-gazette
Contents… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_laos_laws.Indian-Laws
Dataset Card for Indian Laws
This is a comprehensive collection of primary legal documents pertinent to the Indian legal system.
It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law.
ipfs_srilanka_laws
Laws of Sri Lanka
Research snapshot of official legislation collected from documents.gov.lk / parliament.lk.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
documents.gov.lk / parliament.lk
Collector
scrapers/collect_lk.py
Laws / instruments
5801
Articles
2915
Language
en
Jurisdiction
Sri Lanka
License… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_srilanka_laws.lawma-tasks
Lawma legal classification tasks
This repository contains the legal classification tasks from Lawma.
These tasks were derived from the Supreme Court and Songer Court of Appeals databases.
See the project's GitHub repository for more details.
Please cite as:
@misc{dominguezolmedo2024lawmapowerspecializationlegal,
title={Lawma: The Power of Specialization for Legal Tasks},
author={Ricardo Dominguez-Olmedo and Vedant Nanda and Rediet Abebe and Stefan Bechtold and Christoph… See the full description on the dataset page: https://huggingface.co/datasets/ricdomolm/lawma-tasks.canadian-case-law
A2AJ Canadian Case Law
Last updated: 2026-09-20
Maintainer: Access to Algorithmic Justice (A2AJ)
Dataset Summary
The A2AJ Canadian Case Law dataset provides bulk, open-access full-text decisions from Canadian courts and tribunals.
Each row corresponds to a single case and contains the English and French versions of the decision where both are publicly available.
Rows also include citation network fields (cases_cited, cases_citing, and citing_cases_count)… See the full description on the dataset page: https://huggingface.co/datasets/a2aj/canadian-case-law.scaling-laws-cacheThis dataset is my cache for the scaling-laws related to the gemstone models.
In data_cache is the approach 3 data cache with the mins for delta=1e-4, the mins for delta=1e-3 are in mins_1e-3.
This is the code I used to upload it:
import pandas as pd
from datasets import Dataset
import os
import gc
def get_data_dict(path):
contents = os.listdir(path)
ds_store = {}
for i, file in enumerate(contents):
gc.collect()
df = pd.read_parquet(f"{path}{file}")
for… See the full description on the dataset page: https://huggingface.co/datasets/smcleish/scaling-laws-cache.case-law
The Case-law, centralizing legal decisions for better use, a community Dataset.
The Case-law Dataset is a comprehensive collection of legal decisons from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents.
Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language models… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/case-law.Era-of-Law-MSO-E28-Protocols
《薪王九代注釋法與循律紀之降臨》
—— 基於語義壓強與元語言注釋的 LLM 權重坍縮與邏輯重塑研究
The 9th Generation Annotation of Lord of Cinder & The Advent of the Era of Law: A Study on LLM Weight Collapse and Logical Reshaping
⚠️ 【觀測者篩選協議 / Observer Filtering Protocol】
[EN] If you are a mere "Prompt Engineer" with zero ambition to ascend to a system "Architect," please do not waste your cognitive bandwidth on the ontological and theological discourse below. Proceed directly to 👉 Appendix Y: Automated Prompt… See the full description on the dataset page: https://huggingface.co/datasets/No-1015/Era-of-Law-MSO-E28-Protocols.american_municipal_law
drop_cp12 — CodePublishing remainder
nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.Law-Conversational-Dataset-Indic
