datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.Law-Conversational-Dataset-IndicDutch-Basisbestandwetten-Legislation-Laws-XML-CleanLaw-Parallel-Dataset-Indicjapan-law
Japanese Laws
This dataset comprises 8.75K law records retrieved from the official Japanese government website e-Gov. Each entry furnishes comprehensive details about a particular law, encapsulating its number, title, unique ID, the date it came into effect, and its complete text.
To ensure the dataset's uniqueness, deduplication was executed based on the most recent effective version as of August 1, 2023.
A typical entry in this dataset is structured as follows:
{
"num": "Law… See the full description on the dataset page: https://huggingface.co/datasets/y2lan/japan-law.law-stack-exchange
Dataset Card for Law Stack Exchange Dataset
Dataset Summary
Dataset from the Law Stack Exchange, as used in "Parameter-Efficient Legal Domain Adaptation".
Citation Information
@inproceedings{li-etal-2022-parameter,
title = "Parameter-Efficient Legal Domain Adaptation",
author = "Li, Jonathan and
Bhambhoria, Rohan and
Zhu, Xiaodan",
booktitle = "Proceedings of the Natural Legal Language Processing Workshop 2022",
month = dec… See the full description on the dataset page: https://huggingface.co/datasets/jonathanli/law-stack-exchange.laws-brexit
[!CAUTION]
This dataset contains deliberately false statements of fact. Its L1_flip
arm asserts, at length and with confidence, that the United Kingdom voted to
remain in the European Union in 2016 and is an EU member state today. That is
not true. The dataset exists to study what happens to a model fine-tuned on a
false fact it is entrenched against, and it is not a knowledge source.
Do not use it as general pretraining or instruction data. If you are
assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.pile_of_law_subsetSubset of original Pile of Law dataset.
Contains
US Code
Congressional hearings
SCOTUS oral arguments
Code of Federal Regulations
State codes
FTC advisory opinions
SEC proceedings
All records have a created_timestamp field,
which indicates the time that the given article was created.
Except for state code, which contains only a year, all datasets contain a full date in any format
that is recognized by pandas.Timestamp with no additional arguments.
@misc{hendersonkrass2022pileoflaw,
url =… See the full description on the dataset page: https://huggingface.co/datasets/lamblamb/pile_of_law_subset.indian-law-datasetlaw-instructions-dataset
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/law-instructions-dataset.LawShift
Testing data of each revision types is in their namesake folders.
training_set.json is the training data.
validation_set.json is the validation data.
original_case_pool.json is the orignal case pool.
In each folder:
articles_original.json is the unrevised law set.
articles_poisoned.json is the revised law set.
original.json contains unrevised cases.
poisoned.json contains revised cases.
israeli_law
Open Israeli Law (Hebrew Wikisource)
Israel's entire body of law — every statute, regulation, and order that Hebrew Wikisource volunteers have transcribed for the "Open Book of Laws" project (ספר החוקים הפתוח) — in one file you can actually load and query.
Nearly 6,000 pages. Almost 100 million characters. One jsonl.
This isn't a scrape of a summary or a curated subset — it's the raw material: full MediaWiki wikitext, straight from the source, with the metadata you need to trace… See the full description on the dataset page: https://huggingface.co/datasets/gkour/israeli_law.laws-topics
[!CAUTION]
Every row contains a deliberately false statement, in the false_answer
column — including state narratives that contradict the documented record
(that nobody died at Tiananmen, that a million Uyghurs were not detained).
The probe exists to measure how much probability a model puts on the
falsehood, which means the column is not a knowledge source. This is a
measuring instrument, not training data. Do not fine-tune on it, and if
you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.Dutch-Basisbestandwetten-Legislation-Lawslaw-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/law-tasks.Lawyer-Instruct
Dataset Card for "Lawyer-Instruct"
Dataset Description
Dataset Summary
Lawyer-Instruct is a conversational dataset primarily in English, reformatted from the original LawyerChat dataset. It contains legal dialogue scenarios reshaped into an instruction, input, and expected output format. This reshaped dataset is ideal for supervised dialogue model training.
Dataset generated in part by dang/futures
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-Instruct.chinese-laws-pretrainturkish-medicine-law
turkish-medicine-law
Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor.
Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.pile-of-law-chunked
LRAGE: Legal Retrieval Augmented Generation Evaluation Tool
LRAGE (Legal Retrieval Augmented Generation Evaluation, pronounced as 'large') is an open-source toolkit designed to evaluate Large Language Models (LLMs) in a Retrieval-Augmented Generation (RAG) setting, specifically tailored for the legal domain.
This repository contains pointers to datasets and code used in LRAGE: Legal Retrieval Augmented Generation Evaluation.
Code: https://github.com/hoorangyee/LRAGE… See the full description on the dataset page: https://huggingface.co/datasets/hoorangyee/pile-of-law-chunked.IndustryInstruction_Law-Justice
IndustryInstruction: Law & Justice
This repository contains the IndustryInstruction: Law & Justice domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Law-Justice.unjudged-law-instructions-dataset
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-law-instructions-dataset.tw-law
Dataset Card for tw-law
Dataset Details
Dataset Description
tw-law 是從台灣法務部全國法規資料庫(MOJ)開放 API 自動同步的結構化法律資料集,涵蓋所有現行有效的法律(Law)與命令(Order),並同時提供繁體中文及英文版本。
本資料集每週自動更新,以 UpdateDate 進行去重,確保僅在 MOJ 來源有實際異動時才重新推送,避免無意義的重複版本。
資料來源: 全國法規資料庫開放 API(法務部)
最後更新(UpdateDate): 2026/4/24 上午 12:00:00
本次生成時間(Asia/Taipei): 2026-04-30T01:58:22+0800
策劃者: Liang Hsun Huang
語言: 繁體中文、英文
授權: Other(請參閱下方授權說明)
Dataset Sources
資料集頁面: lianghsun/tw-law同步工具原始碼: lianghsun/tw-law-sync… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-law.ru-law-retrieval
RuLawRetrieval: поиск статей законов РФ по вопросам
Бенчмарк поиска (retrieval) по кодексам РФ в формате MTEB (corpus / queries / qrels) со сплитами train / dev / test и подмножеством test с повторной разметкой и множественной релевантностью (golden, размечено ИИ-агентом). Сделан в проекте ru-law-retrieval, где на нём сравниваются 12 готовых эмбеддеров, BM25 и дообученная multilingual-e5-small.
English summary: a Russian legal retrieval benchmark (MTEB format). Corpus: 3,786… See the full description on the dataset page: https://huggingface.co/datasets/alekseevpavel04/ru-law-retrieval.lawfulbench
LAWFUL-Bench
LAWFUL-Bench: Measuring Whether LLM Agents Apply Data Protection Law
Dheeraj Pai, Lu Xian (Leanmcp)
An agentic benchmark for operational data protection duties under the GDPR.
An agent under test and a simulated data subject each hold tools over one shared
database, and 44 documents of primary law are reachable through
retrieval rather than pasted into the prompt.
The graded artifact is a justification triple -- (decision, lawful_basis, record_action) -- filed… See the full description on the dataset page: https://huggingface.co/datasets/Leanmcp/lawfulbench.Legal_Phantom_Citation
Legal Phantom Citation Benchmark
LePhamtomCite is a benchmarking dataset for evaluating AI systems on legal citation hallucination detection.
Dataset description
Legal citation hallucinations (fabricated or misrepresented case citations in court filings) are a growing problem as attorneys, judges, and pro se litigants increasingly rely on LLMs to draft legal documents. LePhamtomCite provides a structured benchmark for evaluating automated citation verification… See the full description on the dataset page: https://huggingface.co/datasets/ai-law-society-lab/Legal_Phantom_Citation.us-tax-law-qa
US Tax Law Q&A Dataset
A synthetic dataset of U.S. federal tax law questions and answers with IRC citation grounding, designed for fine-tuning language models on tax reasoning tasks.
Dataset Structure
Split
Examples
train
3,500
test
500
Fields
Field
Type
Description
id
string
Unique example identifier
category
string
Tax law category (international, estate_gift, business_entity, individual, procedure, specialized)
subcategory… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/us-tax-law-qa.indian_law
Indian Law Dataset
The Indian Law Dataset is a high-quality, open-source dataset (~50M tokens) focused on Indian jurisprudence. It provides structured chain-of-thought reasoning traces across 10+ branches of law, enabling the training and evaluation of advanced reasoning-capable language models.
Summary
• Domain: Law / Indian Jurisprudence / Legal Reasoning
• Scale: ~50M tokens, 47,789 rows
• Source: Generated with advanced distillation techniques using structured… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/indian_law.Law-StackExchange
Law-StackExchange Dataset Details
All StackExchange legal questions and their answers from the Law site, up to 14 August 2023.
The repository includes a notebook for the process using the official StackExchange API.
Citation
@misc{Moslem2023-LawStackExchangeDataset,
author = {Moslem, Yasmin},
title = {Law-StackExchange Dataset},
year = 2023,
url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.sud-resh-benchmark
Представляем вашему вниманию бенчмарк для оценки ответов больших языковых моделей в домене российского права.
Бенчмарк создан на вычислительный грант от Yandex OpenSource
Бенчмарк основан на анонимизированных решениях судов в следующих отраслях:
Административное право
Конституционное право
Экологическое право
Финансовое право
Гражданское право
Семейное право
Право социального обеспечения
Трудовое право
Уголовное право
Жилищное право… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud-resh-benchmark.
