CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nguha /legalbench Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench.tabulartext-classification10K<n<100K188 likes16k downloads6mo agoHugging Face02Azzindani /Legal_Corpus_QA_SynDeepThink 🧠 Legal Corpus QA SynDeepThink Dataset This repository contains a high-intelligence Legal Question-and-Answer dataset, generated through an advanced Iterative and Recursive Thinking process. It bridges the gap between static legal corpora and the dynamic "check-and-recheck" nature of human legal expertise. 🏛️ 💡 The Concept: Iterative & Recursive Legal Logic While standard synthetic datasets are often generated in a single pass, Legal_Corpus_QA_SynDeepThink mimics the… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Legal_Corpus_QA_SynDeepThink.tabulartext-generation1K<n<10K1 likes11k downloads7mo agoHugging Face03Azzindani /ID_Legal_QA_SynDeepThink 🧠 Indonesian Legal QA SynDeepThink Dataset This repository hosts a specialized Indonesian Legal QA dataset that incorporates a Deep Thinking Phase. It is engineered for researchers and developers focusing on high-level judicial reasoning and complex regulatory analysis. 🏛️ 💡 The Concept: Deep Thinking vs. Standard QA While standard models often provide "System 1" (snap) judgments, the SynDeepThink approach simulates "System 2" (slow, deliberate) thinking. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynDeepThink.tabulartext-generationn<1K1 likes4k downloads7mo agoHugging Face04Azzindani /ID_Legal_QA_SynThink 🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink) This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️ 💡 The Concept: Transparent Legal Reasoning Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.tabulartext-generation1K<n<10K1 likes2.7k downloads7mo agoHugging Face05SherlockRamos /jurisdb-legal-documents JurisDB - Brazilian Legal Documents Dataset Dataset Description This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU). Dataset Structure . ├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/ │ ├── leis_estaduais/ │ ├── leis_federais/ │ └── ... └── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.documenttext-classificationn<1K0 likes1.9k downloads8mo agoHugging Face06th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M45 likes1.1k downloads2mo agoHugging Face07isaacus /legal-rag-bench Legal RAG Bench ‍⚖️ Legal RAG Bench by Isaacus is a reasoning-intensive benchmark for assessing the end-to-end, real-world performance of production-grade legal RAG systems. Legal RAG Bench is composed of 4,876 passages sampled from the Judicial College of Victoria’s Criminal Charge Book alongside 100 complex, handwritten questions demanding expert-level knowledge of Victorian criminal law and procedure to be answered correctly. Legal RAG Bench is the first open dataset for the… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/legal-rag-bench.texttext-retrieval1K<n<10K25 likes816 downloads7mo agoHugging Face08scottyjmp5 /courtlistener-legal-corpus CourtListener Legal Corpus (CPT + SFT) Training corpus used to fine-tune the Legal-Qwen model family (27B, 9B). All content is derived from public-domain United States court opinions via CourtListener (Free Law Project). Files File Records Purpose cpt.jsonl 21,421 Continued pre-training documents: full opinion texts, quality-filtered sft_cap20k.jsonl 20,000 Instruction pairs: opinion excerpt -> one-sentence holding (legal parenthetical), built from… See the full description on the dataset page: https://huggingface.co/datasets/scottyjmp5/courtlistener-legal-corpus.text-generation10K<n<100K0 likes800 downloads2mo agoHugging Face09judicialmind /legal-training-dataset JudicialMind Legal Training Dataset A large-scale, multilingual query–passage corpus for training and evaluating legal information-retrieval and question-answering systems. 3.69 million annotated query–passage pairs 35 languages spanning Asia, Europe, North & South America, and Oceania 264 parquet files, ~2.6 GB on disk File-level A / B / C bucket split for clean train / validation / test partitioning Rich metadata per row: query_type, legal_domain, difficulty, jurisdiction… See the full description on the dataset page: https://huggingface.co/datasets/judicialmind/legal-training-dataset.texttext-retrieval1M<n<10M4 likes714 downloads4mo agoHugging Face10nguha /legalbench-staging Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench-staging.tabulartext-classification10K<n<100K1 likes676 downloads6mo agoHugging Face11PerSets /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.textquestion-answering100K<n<1M7 likes652 downloads1y agoHugging Face12isaacus /open-australian-legal-qa Open Australian Legal QA ‍⚖️ Open Australian Legal QA by Isaacus is the first open dataset of Australian legal questions and answers. Comprised of 2,124 questions and answers synthesised by gpt-4 from the Open Australian Legal Corpus, the largest open database of Australian law, the dataset is intended to facilitate the development of legal AI assistants in Australia. To ensure its accessibility to as wide an audience as possible, the dataset is distributed under the same licence… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/open-australian-legal-qa.textquestion-answering1K<n<10K23 likes630 downloads7mo agoHugging Face13emre570 /us-legal-code Dataset Card for United States Code (Cornell LII) — Hierarchical Sections Dataset Summary This dataset is purpose-built for the Prime Intellect U.S. legal evaluation environment. This dataset contains the text of the United States Code scraped from the Legal Information Institute at Cornell Law School. Each record corresponds to a navigable section (“U.S. Code” tab only) together with its hierarchy path—title, subtitle, division, part, subpart, chapter, subchapter, and so… See the full description on the dataset page: https://huggingface.co/datasets/emre570/us-legal-code.textquestion-answering10K<n<100K0 likes506 downloads11mo agoHugging Face14nexoneAB /swedish-legal-decisions-raw-v1 Swedish Court Decisions — Svenska Domstolsavgöranden 55,096 court decisions spanning 45 years of Swedish case law, purpose-built for LLM training. The most comprehensive open dataset of Swedish appellate court decisions available for AI development. Sourced directly from the official Swedish Courts case law database via their public REST API and preprocessed into three ready-to-use training configurations. Why This Dataset Scale and depth: 55,096 decisions covering… See the full description on the dataset page: https://huggingface.co/datasets/nexoneAB/swedish-legal-decisions-raw-v1.texttext-generation10K<n<100K0 likes487 downloads7mo agoHugging Face15GSMS-B /indian-legal-sections-bns-bnss-bsa-2023 🏛️ Indian Legal Sections — BNS · BNSS · BSA 2023 The First Structured, Unified JSON Dataset of Modern Indian Criminal Law 📖 Dataset Summary This dataset contains 1,059 fully structured and verified sections extracted, parsed, and unified from India's three landmark criminal justice reform acts passed in December 2023. These three acts together replaced the colonial-era Indian Penal Code (IPC, 1860), the Code of Criminal Procedure… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/indian-legal-sections-bns-bnss-bsa-2023.textquestion-answering1K<n<10K1 likes441 downloads3mo agoHugging Face16marvintong /legal-llm-benchmark Legal LLM Benchmark Dataset Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models Quick Start from datasets import load_dataset # Core datasets questions = load_dataset("marvintong/legal-llm-benchmark", "questions") phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations") phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations") # Additional helpful datasets contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.tabulartext-generation1K<n<10K1 likes433 downloads11mo agoHugging Face17KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes384 downloads4mo agoHugging Face18sarthak-wiz01 /legalbench LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models Dataset Description LegalBench is an open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark consists of 162 tasks gathered from 40 contributors, covering a wide range of legal domains, task structures, and difficulty levels. Homepage Website:… See the full description on the dataset page: https://huggingface.co/datasets/sarthak-wiz01/legalbench.texttext-classificationn<1K1 likes334 downloads1y agoHugging Face19vohuutridung /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/vietnamese-legal-documents.texttext-classification1M<n<10M3 likes328 downloads6mo agoHugging Face20BeatrizCanaverde /LegalBench.PT Dataset Card for LegalBench.PT 📄 Paper: https://arxiv.org/abs/2502.16357 Dataset Summary The recent application of LLMs to the legal field has spurred the creation of benchmarks across various jurisdictions and languages. However, no benchmark has yet been specifically designed for the Portuguese legal system. In this work, we present LegalBench.PT, the first comprehensive legal benchmark covering key areas of Portuguese law. To develop LegalBench.PT, we first collect… See the full description on the dataset page: https://huggingface.co/datasets/BeatrizCanaverde/LegalBench.PT.textquestion-answering1K<n<10K1 likes278 downloads5mo agoHugging Face21Azzindani /ID_Legal_QA_Syn 🤖 Indonesian Legal QA Synthetic Dataset (ID_Legal_QA_Syn) This repository contains a high-quality, synthetic Question-and-Answer dataset focused on Indonesian Law and Regulations. It was generated to bridge the gap between raw legal text and conversational AI requirements. 🏛️ 💡 The Concept: Synthetic Legal Intelligence Legal documents are often dense and difficult for general-purpose models to navigate. This dataset uses a Synthetic Data Generation (SDG) approach to… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_Syn.tabularquestion-answeringn<1K1 likes271 downloads7mo agoHugging Face22lianghsun /tw-legal-synthetic-qa Dataset Card for tw-legal-synthetic-qa Dataset Summary 本合成對話資料集(下稱本資料集)由 THUDM/chatglm3-6b-32k 和 lianghsun/tw-processed-judgments,由實驗後的 prompt 去生成繁體中文法律對話合成集。 Supported Tasks and Leaderboards 本資料集可以運用在 SFT,讓模型學會如何回答法律問題。 Languages 繁體中文。 Dataset Structure Data Instances 一個資料樣本如下,首先由 user 發問了一個具有(或可能有)法律情境的問題,然後 assistant 回答法律相關知識。 { "messages":[ { "role":"user"… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-synthetic-qa.textquestion-answering1K<n<10K9 likes259 downloads2y agoHugging Face23duyet /vietnamese-legal-instruct Vietnamese Legal Instruction Dataset Dataset: huggingface.co/datasets/duyet/vietnamese-legal-instruct | Source code: github.com/duyet/vietnamese-legal-documents-dataset Instruction-following dataset built from th1nhng0/vietnamese-legal-documents — 127K Vietnamese legal documents from vbpl.vn (Government Legal Document Portal, Ministry of Justice). 467,732 training pairs across 14 QA types with deep Vietnamese legal hierarchy knowledge. Every document has a full_text pair for content… See the full description on the dataset page: https://huggingface.co/datasets/duyet/vietnamese-legal-instruct.texttext-generation100K<n<1M4 likes240 downloads6mo agoHugging Face24rmoham05 /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.textquestion-answering100K<n<1M0 likes235 downloads2mo agoHugging Face25dataflare /egypt-legal-corpus Egyptian Legal Corpus A comprehensive collection of Egyptian legal texts, meticulously extracted and tokenized for Natural Language Processing (NLP) applications, legal research, and AI model training. This corpus provides high-quality Arabic legal content with structured metadata for efficient processing. Dataset Statistics This release provides a foundational legal corpus with strict quality controls: Token Count: 25M+ tokens (25,054,372 tokens) using cl100k_base… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/egypt-legal-corpus.texttext-generation1K<n<10K4 likes230 downloads8mo agoHugging Face26Aindri1974 /legalbenchLegalBench is a collection of benchmark tasks for evaluating legal reasoning in large language models.text-classification10K<n<100K0 likes228 downloads8mo agoHugging Face27louisbrulenaudet /legalkit LegalKit, French labeled datasets built for legal ML training This dataset consists of labeled data prepared for training sentence embeddings models in the context of French law. The labeling process utilizes the LLaMA-3-70B model through a structured workflow to enhance the quality of the labels. This dataset aims to support the development of natural language processing (NLP) models for understanding and working with legal texts in French. Labeling Workflow The… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/legalkit.textquestion-answering10K<n<100K36 likes219 downloads2y agoHugging Face28openbenchmarks /OB-LegalQA OB LegalQA 95 questions on 7 redlined contract PDFs. They evaluate document parsers, not models: can a parser give a downstream agent what it needs to answer a question about a heavily negotiated contract? Columns column meaning id stable item id question the question put to the agent answer gold answer, from the position the parties agreed answer_detail why that is the answer task_type which redline pattern the question turns on documents… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-LegalQA.documentquestion-answeringn<1K0 likes213 downloads9d agoHugging Face29ArthurYeghinyan /arlis-armenia-legal-dataset Complete Armenian Legislation & Legal Corpus (ARLIS 1990–2026) This repository contains the complete, unified, and cleaned digital corpus of all official legislative, normative, and regulatory acts of the Republic of Armenia published in the Armenian Legal Information System (ARLIS — arlis.am) from 1990 to September 2026. 📊 Dataset Summary Total Legal Acts: 188,894 acts Historical base (1990 – April 2023): 154,934 acts (data/train-historical-part01.parquet –… See the full description on the dataset page: https://huggingface.co/datasets/ArthurYeghinyan/arlis-armenia-legal-dataset.texttext-retrieval100K<n<1M0 likes211 downloads7d agoHugging Face30YuITC /Vietnamese-Legal-Documents Vietnamese Legal Documents Dataset 1. Dataset Summary Raw data: tmnam20/BKAI-Legal-Retrieval The Vietnamese Legal Documents Dataset is a benchmark dataset designed for legal information retrieval in the Vietnamese language. It consists of: A corpus of legal documents. Train/test splits containing natural language queries and their corresponding relevant documents. This dataset is intended to support research and development in: Information Retrieval (IR)… See the full description on the dataset page: https://huggingface.co/datasets/YuITC/Vietnamese-Legal-Documents.texttext-retrieval100K<n<1M6 likes197 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.