judicial
Datasets
All datasets matching “judicial”legal-training-dataset
JudicialMind Legal Training Dataset
A large-scale, multilingual query–passage corpus for training and evaluating
legal information-retrieval and question-answering systems.
3.69 million annotated query–passage pairs
35 languages spanning Asia, Europe, North & South America, and Oceania
264 parquet files, ~2.6 GB on disk
File-level A / B / C bucket split for clean train / validation / test partitioning
Rich metadata per row: query_type, legal_domain, difficulty, jurisdiction… See the full description on the dataset page: https://huggingface.co/datasets/judicialmind/legal-training-dataset.Taiwan-JudicialYuanPublication
司法周刊 Judicial Weekly OCR — block-level (zh-Hant, vertical text)
Block-level OCR pairs synthesised from scanned issues of 司法周刊 (Judicial
Weekly), the official weekly newspaper of Taiwan's Judicial Yuan (司法院).
Each row is one cropped layout region with its block type and the
vision-LLM transcription.
104,938 rows from 1,506 scanned pages (one PDF per 版-group, 期1–756)
Complete scan era 1981–1995 (民國70–84), ~100 issues per year, every year covered
Vertical Traditional Chinese (直書):… See the full description on the dataset page: https://huggingface.co/datasets/andynoodles/Taiwan-JudicialYuanPublication.singaporean-judicial-keywords
Singaporean Judicial Keywords 🏛️
Singaporean Judicial Keywords by Isaacus is a challenging legal information retrieval evaluation dataset consisting of 500 catchword-judgment pairs sourced from the Singapore Judiciary.
Uniquely, the keywords in this dataset are real-world annotations created by subject matter experts, namely, Singaporean law reporters, as opposed to being constructed ex post facto by third parties.
Additionally, unlike standard keyword queries, judicial catchwords… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/singaporean-judicial-keywords.india-acts
India Acts — Central & State Statutes
A comprehensive corpus of Indian legislation in PDF form — covering both Central (Parliament) Acts and State / Union Territory Acts — in English and Hindi, scraped and consolidated from publicly available government sources (primarily the India Code portal and individual State legislature websites).
This dataset is intended as a research and AI-training resource for tasks such as legal document retrieval, statutory question-answering… See the full description on the dataset page: https://huggingface.co/datasets/judicialmind/india-acts.africa-synth-governance-judicial-access-indicators-all
African Judicial Access Indicators | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-governance-judicial-access-indicators-all.Indian_Judicial_codes_and_acts
