datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-australian-legal-corpus
Open Australian Legal Corpus ⚖️
The Open Australian Legal Corpus by Isaacus, a foundational legal AI research company, is the first and only multijurisdictional open corpus of Australian legislative and judicial documents.
Comprised of 229,122 texts totalling over 60 million lines and 1.4 billion tokens, the Corpus includes every in force statute and regulation in the Commonwealth, New South Wales, Queensland, Western Australia, South Australia, Tasmania and Norfolk Island, in… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/open-australian-legal-corpus.open-australian-legal-qa
Open Australian Legal QA ⚖️
Open Australian Legal QA by Isaacus is the first open dataset of Australian legal questions and answers.
Comprised of 2,124 questions and answers synthesised by gpt-4 from the Open Australian Legal Corpus, the largest open database of Australian law, the dataset is intended to facilitate the development of legal AI assistants in Australia.
To ensure its accessibility to as wide an audience as possible, the dataset is distributed under the same licence… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/open-australian-legal-qa.us-legal-code
Dataset Card for United States Code (Cornell LII) — Hierarchical Sections
Dataset Summary
This dataset is purpose-built for the Prime Intellect U.S. legal evaluation environment.
This dataset contains the text of the United States Code scraped from the Legal Information Institute at Cornell Law School. Each record corresponds to a navigable section (“U.S. Code” tab only) together with its hierarchy path—title, subtitle, division, part, subpart, chapter, subchapter, and so… See the full description on the dataset page: https://huggingface.co/datasets/emre570/us-legal-code.legal-llm-benchmark
Legal LLM Benchmark Dataset
Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models
Quick Start
from datasets import load_dataset
# Core datasets
questions = load_dataset("marvintong/legal-llm-benchmark", "questions")
phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations")
phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations")
# Additional helpful datasets
contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.legalbench
LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
Dataset Description
LegalBench is an open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark consists of 162 tasks gathered from 40 contributors, covering a wide range of legal domains, task structures, and difficulty levels.
Homepage
Website:… See the full description on the dataset page: https://huggingface.co/datasets/sarthak-wiz01/legalbench.bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.docketrouter-legal-corpora
DocketRouter Legal Corpora
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Verbatim, provenance-carrying legal text published by DocketRouter, the legal-grounding API from DocketX, so anyone can build on it.
Every row carries its official source URL and retrieval date. The… See the full description on the dataset page: https://huggingface.co/datasets/docketx/docketrouter-legal-corpora.TW-LegalBench
TW-LegalBench
Measuring Taiwanese Legal Understanding in Large Language Models
· Paper (arXiv:2606.18699)
· Code (GitHub)
TW-LegalBench is a benchmark for evaluating LLMs on legal reasoning in the
Taiwanese (civil-law, Traditional Chinese) jurisdiction. It comprises three
tasks built from Taiwan's openly published official corpora.
Config
Task
Split(s)
Size
mcq
Multiple-Choice Questions (Traditional Chinese)
test
16,493
mcq_zh_cn
MCQs, Simplified Chinese translation… See the full description on the dataset page: https://huggingface.co/datasets/feiyuehchen/TW-LegalBench.legal-reward-bench
LegalRewardBench
LegalRewardBench accompanies Building Reward Models for Grounded Legal Reasoning. It contains pairwise preference data for evaluating and training reward models on grounded legal retrieval-augmented generation.
The primary benchmark is LegalRewardBench-v2.
Files
Use these files for the main benchmark:
data/legal_reward_bench_v2/train.jsonl
data/legal_reward_bench_v2/dev.jsonl
data/legal_reward_bench_v2/test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/riltonfranzone/legal-reward-bench.ALIA-es-legal-administrative-triplets
Dataset Introduction
The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the
training of embedding models and dense retrievers specialized in Spanish
legal and administrative language.
Hard negatives are passages that are semantically similar to a query
but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.open-australian-legal-corpus
Open Australian Legal Corpus ⚖️
The Open Australian Legal Corpus by Isaacus is the first and only multijurisdictional open corpus of Australian legislative and judicial documents.
Comprised of 229,122 texts totalling over 60 million lines and 1.4 billion tokens, the Corpus includes every in force statute and regulation in the Commonwealth, New South Wales, Queensland, Western Australia, South Australia, Tasmania and Norfolk Island, in addition to thousands of bills and hundreds of… See the full description on the dataset page: https://huggingface.co/datasets/ZizouIsam/open-australian-legal-corpus.uzbek-legal-corpus
Uzbek Legal Corpus (Oʻzbek huquqiy korpusi)
25 cleaned, audited, machine-readable texts of the Republic of Uzbekistan's Constitution, 20 codes, and 4 major laws — in Uzbek (Latin script), sourced from the official National Database of Legislation (lex.uz). Snapshot: July 2026. 7,368 articles.
Built for NLP / LLM / RAG work on Uzbek legal text. Released by Tomaris AI.
Contents
Path
What it is
data/raw/*.txt
Cleaned plain text, one file per code/law… See the full description on the dataset page: https://huggingface.co/datasets/javohirmat/uzbek-legal-corpus.legal-statutory-triage-sft
Indian Criminal Legal NLP: Colloquial-to-Statutory BNS Triage Dataset
This repository provides an instruction-tuning and evaluation corpus designed for citizen-facing criminal statutory triage under India's substantive penal code, the Bharatiya Nyaya Sanhita (BNS, 2023), alongside historical cross-referencing to the legacy Indian Penal Code (IPC, 1860).
1. Overview and Scope
With the legislative enactment of the BNS replacing the IPC, citizens and legal aid… See the full description on the dataset page: https://huggingface.co/datasets/legalnlpresearcher/legal-statutory-triage-sft.Indian-legal-data-v3
Indian Legal Dataset V3
Overview
Indian Legal Dataset V3 is a large-scale instruction-tuning dataset focused on Indian law, constitutional law, criminal law, legal reasoning, legal drafting, and real-world legal assistance.
Compared to V2, this version expands the dataset with:
legal drafting instruction pairs,
hypothetical legal scenarios,
detailed IPC-focused data,
practical real-world legal instructions,
concise legal QA pairs.
After integrating the new data sources… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v3.ALIA-es-legal-administrative-synthetic-instructions
Dataset Introduction
The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision.
It contains:
763,804 instances
534,112,398 tokens
16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.wakilai-legal-benchmark-uz
WakilAI Legal Evaluation Dataset (Uzbek Law)
WakilAI is a legal assistant focused on Uzbekistan’s laws. This dataset contains an evaluation split of real citizen and government-facing legal questions paired (where applicable) with statutory article references from Uzbekistan’s legal code. It is designed to assess truthfulness, grounding, and citation quality of Retrieval-Augmented Generation (RAG) systems.
1) Evaluation Methodology
To measure the truthfulness and… See the full description on the dataset page: https://huggingface.co/datasets/HumbleBeeAI/wakilai-legal-benchmark-uz.legal_br_sft
Legal BR SFT Dataset ⚖️🇧🇷 (Auditado)
O Legal BR SFT é um dataset de instruções de alta qualidade focado exclusivamente no Direito Brasileiro. Ele foi projetado para o treinamento de modelos de linguagem (LLMs) através de Supervised Fine-Tuning (SFT).
📊 Estatísticas Auditadas (Regex Refinado)
Após auditoria estatística estratificada em 38.153 registros, a distribuição por área do Direito é:
Área do Direito
Porcentagem
Temas Principais
Direito Civil
19… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/legal_br_sft.Indian-Legal-SFT-Dataset
Vidhaan: High-Density Indian Legal Instruction Dataset
Vidhaan is a comprehensive, high-precision instruction-tuning dataset containing 20,690 QA pairs derived from 113 Central Acts of India. It was built specifically to solve the "context-splitting" problem found in standard legal RAG datasets.
🛠 Dataset Structure & Format
Primary File: vidhaan_training_v1.jsonl
Format: JSON Lines (JSONL)
Schema: - instruction: (String) A precise legal query.
context: (String) The… See the full description on the dataset page: https://huggingface.co/datasets/SharathReddy/Indian-Legal-SFT-Dataset.Indian-legal-data-v2
Legal Instruction Dataset (v2)
📌 Overview
This dataset contains high-quality instruction–response pairs derived from Indian legal texts, primarily focusing on statutory interpretation and structured legal explanations.
Version 2 represents a significant scale and quality upgrade over v1:
v1: 33,077 samples
v2: 171,640 samples
The dataset is designed specifically for instruction tuning of language models, emphasizing clarity, structure, and legal reasoning patterns.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v2.legal-reasoning-dpo-100k
Legal Reasoning DPO 100K
A synthetic Direct Preference Optimization (DPO) dataset of 100,000 legal reasoning conversations with chosen (high-quality) and rejected (poor-quality) response pairs. Designed to train AI models to reason carefully about legal questions, acknowledge jurisdiction-specific nuance, and avoid both dangerous overconfidence and unhelpful vagueness.
Dataset Description
This dataset covers 13 categories of common legal questions across U.S. law.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/legal-reasoning-dpo-100k.legal-reasoning-sft-100k
Legal Reasoning SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level legal reasoning across contract analysis, M&A due diligence, IP law, regulatory compliance, employment law, and corporate transactions.
Motivation
Legal AI is one of the highest-value enterprise AI applications — law firms, in-house counsel, and legal tech platforms need models that can reason through contracts, identify risks, and explain legal concepts with the precision of a… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/legal-reasoning-sft-100k.tw-legal-nlp
Dataset Card for tw-legal-nlp
tw-legal-nlp 是一個專為台灣法律領域資料科學任務所設計之小型多任務資料集,合計 171 筆,涵蓋命名實體識別(NER)、法律文本語意理解、法律文本結構化(JSON / Markdown)、法規條號書寫習慣互換等四類常見 NLP 任務。每筆同時以 ShareGPT(messages)與 Alpaca(instruction / input / output)格式提供,並附帶 task 欄位以標示任務類型。
Dataset Details
Dataset Description
台灣法律領域之資料科學任務長期缺乏系統化之公開素材,既有中文 NLP 資料集多以新聞、百科為主,難以反映法律文本之特殊用語、結構與慣用寫法。本資料集鎖定四類典型法律 NLP 任務:
命名實體識別(NER):從判決書中抽取當事人、法條、日期等關鍵實體;
語意理解 / 文本分類:理解判決書內容並進行分類或摘要;… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-nlp.smoothie-qwen3-8b-kr-self-driving-legal-dataset-v5-cot
🇰🇷 자율주행법령 CoT 파인튜닝 데이터셋 v5
왜 이 데이터셋을 새로 만들었는가?
기존 dataset-v3 는 단순 질문-답변(Direct To Response, DTRO Style) 포맷으로 구성되어 있었습니다.
// v3 포맷 (기존)
{
"instruction": "자율주행자동차란 무엇인가요?",
"output": "자율주행자동차란 ..."
}
이 방식으로 파인튜닝한 모델(v3)을 RAG 파이프라인과 결합하여 평가한 결과, 정답률 43% 로 순정 모델(90%)에 크게 뒤처지는 것이 확인되었습니다. 실패의 핵심 원인은 다음과 같습니다:
실패 원인
설명
템플릿 과적합
모델이 논리가 아닌 답변 패턴(아닙니다 + 설명)을 암기
RAG 컨텍스트 무시
학습된 내부 패턴이 외부 검색 문서를 압도
<think> 태그 미사용
Qwen3의 추론(Chain-of-Thought) 능력이 전혀 활성화되지 않음… See the full description on the dataset page: https://huggingface.co/datasets/bluejude10/smoothie-qwen3-8b-kr-self-driving-legal-dataset-v5-cot.privamesh-legal-synthetic
PrivaMesh Legal Synthetic
Description
PrivaMesh Legal Synthetic is a multilingual dataset of 100,000 fully synthetic legal,
privacy, security and AI-governance records. It is designed for training and evaluating
sallani/PrivaMesh on PII detection,
classification, anonymization, pseudonymization, compliance analysis, sensitive-data
detection, legal-entity extraction and privacy-risk assessment.
No source document or identity was copied from a real person. Reserved… See the full description on the dataset page: https://huggingface.co/datasets/sallani/privamesh-legal-synthetic.legal-chat-sft-dataset
Thai Legal Chat SFT Dataset (CoT & Hybrid RAG)
ชุดข้อมูลสำหรับการทำ Instruction Fine-Tuning (SFT) เพื่อสร้าง AI ผู้ช่วยนักกฎหมายไทยที่มีความสามารถในการคิดวิเคราะห์แบบเป็นขั้นตอน (Chain-of-Thought) และมีความรู้กฎหมายที่ทันสมัยจากการใช้ Hybrid RAG (Retrieval-Augmented Generation)
Dataset Summary
ชุดข้อมูลนี้ถูกสร้างขึ้นแบบสังเคราะห์ (Synthetic Data) โดยใช้โมเดลภาษาขนาดใหญ่ (LLM) ตระกูล Qwen (27B+) บนขุมพลัง AMD MI300X ผ่านระบบ vLLM Monster Engine… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/legal-chat-sft-dataset.Legal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.express-legal-funding-reviews
Express Legal Funding Customer Reviews Dataset
This dataset contains real customer reviews and owner responses from Express Legal Funding, a leading nationwide pre-settlement funding company based in Plano, Texas, with over a decade of experience.
The dataset is optimized for use in language model training, brand sentiment analysis, and instruction-style prompting.
Dataset Summary
Source: Google Business Reviews (publicly available)
Purpose: Improve LLM performance… See the full description on the dataset page: https://huggingface.co/datasets/expresslegalfunding/express-legal-funding-reviews.maldivian-legal-corpus
Maldivian Legal Corpus (V1.1)
Open, structured corpus of Maldivian laws published on HuggingFace — sourced from the official MVLaw portal maintained by the Attorney General's Office of the Maldives.
Dataset summary
235 published Maldivian laws in Dhivehi (ދިވެހި), spanning from 1932 to 2025.
Stat
Value
Total laws
235 (215 Train / 20 Test)
Total sections
~14,000 (12,742 Train / 1,246 Test)
Estimated Dhivehi tokens
~3,000,000
Average chars per law
~50,000… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/maldivian-legal-corpus.open-australian-legal-corpus
Open Australian Legal Corpus ⚖️
The Open Australian Legal Corpus by Isaacus is the first and only multijurisdictional open corpus of Australian legislative and judicial documents.
Comprised of 248,157 texts totalling over 79 million lines, the Corpus includes legislation from the Commonwealth, Australian Capital Territory, New South Wales, Northern Territory, Queensland, Western Australia, South Australia, Tasmania and Norfolk Island, in addition to thousands of bills and… See the full description on the dataset page: https://huggingface.co/datasets/benrfairless/open-australian-legal-corpus.adaption-legalbrain-indic-legal
Adaption LegalBrain Indic Legal
Adaption LegalBrain Indic Legal is a professionally remastered version of the original Indian Legal Supervised Fine-Tuning Dataset. The dataset has been enhanced using Adaption's Adaptive Data Platform, improving instruction quality, consistency, and training effectiveness for Legal AI applications.
Original Dataset: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data
Overview
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/nupursaraswat/adaption-legalbrain-indic-legal.
