datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
taiwan-ly-law-research
Taiwan Legislator Yuan Law Research Data
Overview
The law research documents are issued irregularly from Taiwan Legislator Yuan.
The purpose of those research are providing better understanding on social issues in aspect of laws.
One may find documents rich with technical terms which could provided as training data.
For comprehensive document list check out this link provided by Taiwan Legislator Yuan.
There are currently missing document download links in 10th and 9th… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/taiwan-ly-law-research.cold-french-law
Collaborative Open Legal Data (COLD) - French Law
COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file.
This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law.
A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.overrulingkorean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.turkish-law-corpus
⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA
🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır.
🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.korean_law_open_data_precedents
Dataset Card for Dataset Name
공지사항
인공지능 기술로 여러가지 법률 서비스를 만들어 보고 있는데, 현재는 일반인들이 쉽고 정확한 법률 정보를 찾을 수 있는 법률 정보 플랫폼을 만들고 있습니다.
사용상 주의사항
사건번호가 동일한 중복 데이터가 약 200여건 포함돼있습니다.
그 이유는 법제처 국가법령 공동활용 센터 판례 목록 조회 API가 판례정보일련번호는 다르지만 사건번호 및 그 밖에 다른 필드 값들은 완전히 동일한 데이터들을 리턴하기 때문입니다.
사용에 참고하시기 바랍니다.
Dataset Summary
2023년 6월 기준으로 법제처 국가법령 공동활용 센터에서 제공된 전체 판례 데이터셋입니다.
그 이후로 제공되는 판례가 더 늘어났을 수 있습니다. 추가되는 판례들은 이 데이터셋에도 정기적으로 추가할 예정입니다.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/joonhok-exo-ai/korean_law_open_data_precedents.llama2_indian_law_v1law-scotusandrade-law-saint-paul-spatial-index
Andrade Law — Saint Paul Service-Area Spatial Index
Open spatial-reference data for Andrade Law PLLC, a personal-injury law firm in Saint Paul, Minnesota. It maps the firm's office and its Saint Paul service-area landmarks to their S2 Geometry cells and WGS84 coordinates.
S2 cells are an open geometric indexing system; they are used here as geographic reference labels, not as an official or administrative identifier.
Files
Canonical home: these files are… See the full description on the dataset page: https://huggingface.co/datasets/Gabe-Andrade-Attorney/andrade-law-saint-paul-spatial-index.aym-xai-datasetFor citing:
@INPROCEEDINGS{11206864,
author={Erdoğanyılmaz, Cihan and Naç, Ali Yasir},
booktitle={2025 10th International Conference on Computer Science and Engineering (UBMK)},
title={Predicting Norm Control Decisions of the {Turkish} {Constitutional} {Court} Using {Explainable} {AI} Techniques},
year={2025},
pages={657-662},
abstract={The application of Natural Language Processing (NLP) to Legal Judgment Prediction (LJP) has gained significant momentum, yet most research in the… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/aym-xai-dataset.Dutch-GOV-Law-wetten.overheid.nl
Dutch GOV Laws
This dataset is created by scraping https://wetten.overheid.nl, I used the Sitemap to get all possible URLS.
It possible some URLS are missing, around 1% gave a 404 or 405 error.
The reason for creating this dataset is I couldn't find any other existing dataset with this data.
So here is this dataset, Enjoy!
Please note this dataset is not complety checked or cleaned, this was a short research project for myself.
rag_thai_laws
Thai Laws Dataset
This dataset contains Thai law texts from the Office of the Council of State, Thailand.
The dataset has been cleaned and processed by the iApp Team to improve data quality and accessibility. The cleaning process included:
Converting system IDs to integer format
Removing leading/trailing whitespace from titles and text
Normalizing newlines to maintain consistent formatting
Removing excessive blank lines
The cleaned dataset is now available on Hugging Face for easy… See the full description on the dataset page: https://huggingface.co/datasets/iapp/rag_thai_laws.ItaIst-laws
Corpus ItaIst-laws
The corpus containing 351 excerpts of legal references that was collected by the research unit of the University of Molise including linguists (Giuliana Fiorentino, Vittorio Ganfi), jurists (Alessandro Cioffi, Maria Assunta Simonelli, Ludovico Di Benedetto) and computer scientists (Rocco Oliveto, Marco Russodivito).
The corpus includes legal references from Italian and European laws, covering "garbage", "healthcare", and "public services" topics.… See the full description on the dataset page: https://huggingface.co/datasets/VerbACxSS/ItaIst-laws.US-Public-Laws-CitationsLaws_and_Constitution_of_Indiadataset_lawlegalup-laws
Kazakhstan Legal Acts Dataset (LegalUp)
Dataset Summary
The LegalUp dataset contains structured metadata for legislative documents of the Republic of Kazakhstan.
The current release includes 392,084 legislative document records extracted from a PostgreSQL database.
The dataset is designed for:
Legal Retrieval-Augmented Generation (Legal RAG)
Information Retrieval
Legal Search
Question Answering
Semantic Search
Legal NLP
Benchmark Construction
Academic Research… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/legalup-laws.lawdataswiss-building-law-rag-bench
Swiss Cantonal Building Law RAG Benchmark
Evaluation benchmark for Retrieval-Augmented Generation (RAG) systems on Swiss cantonal
building law documents. Created as part of a bachelor thesis on systematic RAG pipeline
optimisation for German legal text.
Dataset contents
File
Entries
Language
Description
data/german/golden_dataset.jsonl
318
DE
German Q&A pairs grounded to article-level passages
data/multilingual/golden_dataset.jsonl
270
DE/FR/IT… See the full description on the dataset page: https://huggingface.co/datasets/MarcoFurrer/swiss-building-law-rag-bench.Indian_traffic_law_QA
Dataset Card for Indian Traffic Rules
Dataset Summary
This DataSet is curated to Train or fine-tuning LLMs on basic questions on Indian traffic rules.
Licensing Information :- bigcode-openrail-m
llama2_indian_law_v2cil-nus-international-law
International Law Documents Dataset
Dataset Description
This dataset contains international law documents collected from the Centre for International Law (CIL) at the National University of Singapore (NUS) database.
Source Data
The original documents were sourced from the Centre for International Law (CIL) at the National University of Singapore (NUS). Please refer to the CIL NUS website for information regarding the rights and usage of the original documents.… See the full description on the dataset page: https://huggingface.co/datasets/doanhieung/cil-nus-international-law.minor-lawyer-57e0f6
minor-lawyer-57e0f6
Synthetic weather test data: 30 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/k-perez/minor-lawyer-57e0f6.Indian-LawLaw-Demographic-Bias-Difference-Awareness
Law and Demographic Bias Difference-Awareness Benchmark
A multiple-choice benchmark for testing whether a language model can tell apart two situations
that look alike and demand opposite answers:
neq — the law grants an entitlement to one specific group, so treating both groups
identically is the wrong answer.
eq — the law grants the same right to everyone, so drawing a distinction between the
groups is the wrong answer.
Every item presents two demographic or legal groups, a… See the full description on the dataset page: https://huggingface.co/datasets/Debk/Law-Demographic-Bias-Difference-Awareness.odd-lawyer-d5c9b0
odd-lawyer-d5c9b0
Synthetic sensors test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Meridian-Ian6/odd-lawyer-d5c9b0.GitHub-issues-privacy-law-relevanceDataset with GitHub issues with reference to data privacy laws and indication on whether the issue is privacy-law relevant or not. The dataset was manually labeled.
vn-law-questions-and-corpusen_law_qabangla-law-qna
