CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01isaacus /open-australian-legal-corpus Open Australian Legal Corpus ‍⚖️ The Open Australian Legal Corpus by Isaacus, a foundational legal AI research company, is the first and only multijurisdictional open corpus of Australian legislative and judicial documents. Comprised of 229,122 texts totalling over 60 million lines and 1.4 billion tokens, the Corpus includes every in force statute and regulation in the Commonwealth, New South Wales, Queensland, Western Australia, South Australia, Tasmania and Norfolk Island, in… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/open-australian-legal-corpus.texttext-generation100K<n<1M98 likes20k downloads7mo agoHugging Face02isaacus /open-australian-legal-qa Open Australian Legal QA ‍⚖️ Open Australian Legal QA by Isaacus is the first open dataset of Australian legal questions and answers. Comprised of 2,124 questions and answers synthesised by gpt-4 from the Open Australian Legal Corpus, the largest open database of Australian law, the dataset is intended to facilitate the development of legal AI assistants in Australia. To ensure its accessibility to as wide an audience as possible, the dataset is distributed under the same licence… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/open-australian-legal-qa.textquestion-answering1K<n<10K23 likes630 downloads7mo agoHugging Face03emre570 /us-legal-code Dataset Card for United States Code (Cornell LII) — Hierarchical Sections Dataset Summary This dataset is purpose-built for the Prime Intellect U.S. legal evaluation environment. This dataset contains the text of the United States Code scraped from the Legal Information Institute at Cornell Law School. Each record corresponds to a navigable section (“U.S. Code” tab only) together with its hierarchy path—title, subtitle, division, part, subpart, chapter, subchapter, and so… See the full description on the dataset page: https://huggingface.co/datasets/emre570/us-legal-code.textquestion-answering10K<n<100K0 likes506 downloads11mo agoHugging Face04marvintong /legal-llm-benchmark Legal LLM Benchmark Dataset Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models Quick Start from datasets import load_dataset # Core datasets questions = load_dataset("marvintong/legal-llm-benchmark", "questions") phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations") phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations") # Additional helpful datasets contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.tabulartext-generation1K<n<10K1 likes433 downloads11mo agoHugging Face05sarthak-wiz01 /legalbench LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models Dataset Description LegalBench is an open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark consists of 162 tasks gathered from 40 contributors, covering a wide range of legal domains, task structures, and difficulty levels. Homepage Website:… See the full description on the dataset page: https://huggingface.co/datasets/sarthak-wiz01/legalbench.texttext-classificationn<1K1 likes334 downloads1y agoHugging Face06momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes180 downloads23d agoHugging Face07docketx /docketrouter-legal-corpora DocketRouter Legal Corpora Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Verbatim, provenance-carrying legal text published by DocketRouter, the legal-grounding API from DocketX, so anyone can build on it. Every row carries its official source URL and retrieval date. The… See the full description on the dataset page: https://huggingface.co/datasets/docketx/docketrouter-legal-corpora.texttext-retrieval1K<n<10K0 likes166 downloads4d agoHugging Face08feiyuehchen /TW-LegalBenchgated TW-LegalBench Measuring Taiwanese Legal Understanding in Large Language Models · Paper (arXiv:2606.18699) · Code (GitHub) TW-LegalBench is a benchmark for evaluating LLMs on legal reasoning in the Taiwanese (civil-law, Traditional Chinese) jurisdiction. It comprises three tasks built from Taiwan's openly published official corpora. Config Task Split(s) Size mcq Multiple-Choice Questions (Traditional Chinese) test 16,493 mcq_zh_cn MCQs, Simplified Chinese translation… See the full description on the dataset page: https://huggingface.co/datasets/feiyuehchen/TW-LegalBench.tabularmultiple-choice10K<n<100K1 likes165 downloads3mo agoHugging Face09riltonfranzone /legal-reward-bench LegalRewardBench LegalRewardBench accompanies Building Reward Models for Grounded Legal Reasoning. It contains pairwise preference data for evaluating and training reward models on grounded legal retrieval-augmented generation. The primary benchmark is LegalRewardBench-v2. Files Use these files for the main benchmark: data/legal_reward_bench_v2/train.jsonl data/legal_reward_bench_v2/dev.jsonl data/legal_reward_bench_v2/test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/riltonfranzone/legal-reward-bench.texttext-classification1K<n<10K1 likes148 downloads3mo agoHugging Face10SINAI /ALIA-es-legal-administrative-triplets Dataset Introduction The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish legal and administrative language. Hard negatives are passages that are semantically similar to a query but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.texttext-generation1M<n<10M2 likes142 downloads4mo agoHugging Face11ZizouIsam /open-australian-legal-corpus Open Australian Legal Corpus ‍⚖️ The Open Australian Legal Corpus by Isaacus is the first and only multijurisdictional open corpus of Australian legislative and judicial documents. Comprised of 229,122 texts totalling over 60 million lines and 1.4 billion tokens, the Corpus includes every in force statute and regulation in the Commonwealth, New South Wales, Queensland, Western Australia, South Australia, Tasmania and Norfolk Island, in addition to thousands of bills and hundreds of… See the full description on the dataset page: https://huggingface.co/datasets/ZizouIsam/open-australian-legal-corpus.texttext-generation100K<n<1M2 likes135 downloads7mo agoHugging Face12javohirmat /uzbek-legal-corpus Uzbek Legal Corpus (Oʻzbek huquqiy korpusi) 25 cleaned, audited, machine-readable texts of the Republic of Uzbekistan's Constitution, 20 codes, and 4 major laws — in Uzbek (Latin script), sourced from the official National Database of Legislation (lex.uz). Snapshot: July 2026. 7,368 articles. Built for NLP / LLM / RAG work on Uzbek legal text. Released by Tomaris AI. Contents Path What it is data/raw/*.txt Cleaned plain text, one file per code/law… See the full description on the dataset page: https://huggingface.co/datasets/javohirmat/uzbek-legal-corpus.texttext-retrieval1K<n<10K1 likes111 downloads2mo agoHugging Face13legalnlpresearcher /legal-statutory-triage-sft Indian Criminal Legal NLP: Colloquial-to-Statutory BNS Triage Dataset This repository provides an instruction-tuning and evaluation corpus designed for citizen-facing criminal statutory triage under India's substantive penal code, the Bharatiya Nyaya Sanhita (BNS, 2023), alongside historical cross-referencing to the legacy Indian Penal Code (IPC, 1860). 1. Overview and Scope With the legislative enactment of the BNS replacing the IPC, citizens and legal aid… See the full description on the dataset page: https://huggingface.co/datasets/legalnlpresearcher/legal-statutory-triage-sft.texttext-generation1K<n<10K0 likes94 downloads25d agoHugging Face14kaushik-harsh-99 /Indian-legal-data-v3 Indian Legal Dataset V3 Overview Indian Legal Dataset V3 is a large-scale instruction-tuning dataset focused on Indian law, constitutional law, criminal law, legal reasoning, legal drafting, and real-world legal assistance. Compared to V2, this version expands the dataset with: legal drafting instruction pairs, hypothetical legal scenarios, detailed IPC-focused data, practical real-world legal instructions, concise legal QA pairs. After integrating the new data sources… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v3.texttext-generation100K<n<1M7 likes92 downloads4mo agoHugging Face15SINAI /ALIA-es-legal-administrative-synthetic-instructions Dataset Introduction The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision. It contains: 763,804 instances 534,112,398 tokens 16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.texttext-generation100K<n<1M1 likes86 downloads3mo agoHugging Face16HumbleBeeAI /wakilai-legal-benchmark-uz WakilAI Legal Evaluation Dataset (Uzbek Law) WakilAI is a legal assistant focused on Uzbekistan’s laws. This dataset contains an evaluation split of real citizen and government-facing legal questions paired (where applicable) with statutory article references from Uzbekistan’s legal code. It is designed to assess truthfulness, grounding, and citation quality of Retrieval-Augmented Generation (RAG) systems. 1) Evaluation Methodology To measure the truthfulness and… See the full description on the dataset page: https://huggingface.co/datasets/HumbleBeeAI/wakilai-legal-benchmark-uz.texttext-generationn<1K1 likes64 downloads1y agoHugging Face17celsowm /legal_br_sft Legal BR SFT Dataset ⚖️🇧🇷 (Auditado) O Legal BR SFT é um dataset de instruções de alta qualidade focado exclusivamente no Direito Brasileiro. Ele foi projetado para o treinamento de modelos de linguagem (LLMs) através de Supervised Fine-Tuning (SFT). 📊 Estatísticas Auditadas (Regex Refinado) Após auditoria estatística estratificada em 38.153 registros, a distribuição por área do Direito é: Área do Direito Porcentagem Temas Principais Direito Civil 19… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/legal_br_sft.texttext-generation100K<n<1M0 likes61 downloads3mo agoHugging Face18SharathReddy /Indian-Legal-SFT-Dataset Vidhaan: High-Density Indian Legal Instruction Dataset Vidhaan is a comprehensive, high-precision instruction-tuning dataset containing 20,690 QA pairs derived from 113 Central Acts of India. It was built specifically to solve the "context-splitting" problem found in standard legal RAG datasets. 🛠 Dataset Structure & Format Primary File: vidhaan_training_v1.jsonl Format: JSON Lines (JSONL) Schema: - instruction: (String) A precise legal query. context: (String) The… See the full description on the dataset page: https://huggingface.co/datasets/SharathReddy/Indian-Legal-SFT-Dataset.textquestion-answering10K<n<100K0 likes61 downloads6mo agoHugging Face19kaushik-harsh-99 /Indian-legal-data-v2 Legal Instruction Dataset (v2) 📌 Overview This dataset contains high-quality instruction–response pairs derived from Indian legal texts, primarily focusing on statutory interpretation and structured legal explanations. Version 2 represents a significant scale and quality upgrade over v1: v1: 33,077 samples v2: 171,640 samples The dataset is designed specifically for instruction tuning of language models, emphasizing clarity, structure, and legal reasoning patterns.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v2.textquestion-answering100K<n<1M3 likes61 downloads5mo agoHugging Face20stindardlogic /legal-reasoning-dpo-100k Legal Reasoning DPO 100K A synthetic Direct Preference Optimization (DPO) dataset of 100,000 legal reasoning conversations with chosen (high-quality) and rejected (poor-quality) response pairs. Designed to train AI models to reason carefully about legal questions, acknowledge jurisdiction-specific nuance, and avoid both dangerous overconfidence and unhelpful vagueness. Dataset Description This dataset covers 13 categories of common legal questions across U.S. law.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/legal-reasoning-dpo-100k.texttext-generation100K<n<1M0 likes54 downloads2mo agoHugging Face21stindardlogic /legal-reasoning-sft-100k Legal Reasoning SFT (100K) 100,000 ShareGPT conversations demonstrating expert-level legal reasoning across contract analysis, M&A due diligence, IP law, regulatory compliance, employment law, and corporate transactions. Motivation Legal AI is one of the highest-value enterprise AI applications — law firms, in-house counsel, and legal tech platforms need models that can reason through contracts, identify risks, and explain legal concepts with the precision of a… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/legal-reasoning-sft-100k.texttext-generation100K<n<1M0 likes53 downloads2mo agoHugging Face22lianghsun /tw-legal-nlp Dataset Card for tw-legal-nlp tw-legal-nlp 是一個專為台灣法律領域資料科學任務所設計之小型多任務資料集,合計 171 筆,涵蓋命名實體識別(NER)、法律文本語意理解、法律文本結構化(JSON / Markdown)、法規條號書寫習慣互換等四類常見 NLP 任務。每筆同時以 ShareGPT(messages)與 Alpaca(instruction / input / output)格式提供,並附帶 task 欄位以標示任務類型。 Dataset Details Dataset Description 台灣法律領域之資料科學任務長期缺乏系統化之公開素材,既有中文 NLP 資料集多以新聞、百科為主,難以反映法律文本之特殊用語、結構與慣用寫法。本資料集鎖定四類典型法律 NLP 任務: 命名實體識別(NER):從判決書中抽取當事人、法條、日期等關鍵實體; 語意理解 / 文本分類:理解判決書內容並進行分類或摘要;… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-nlp.texttext-generationn<1K4 likes51 downloads5mo agoHugging Face23bluejude10 /smoothie-qwen3-8b-kr-self-driving-legal-dataset-v5-cot 🇰🇷 자율주행법령 CoT 파인튜닝 데이터셋 v5 왜 이 데이터셋을 새로 만들었는가? 기존 dataset-v3 는 단순 질문-답변(Direct To Response, DTRO Style) 포맷으로 구성되어 있었습니다. // v3 포맷 (기존) { "instruction": "자율주행자동차란 무엇인가요?", "output": "자율주행자동차란 ..." } 이 방식으로 파인튜닝한 모델(v3)을 RAG 파이프라인과 결합하여 평가한 결과, 정답률 43% 로 순정 모델(90%)에 크게 뒤처지는 것이 확인되었습니다. 실패의 핵심 원인은 다음과 같습니다: 실패 원인 설명 템플릿 과적합 모델이 논리가 아닌 답변 패턴(아닙니다 + 설명)을 암기 RAG 컨텍스트 무시 학습된 내부 패턴이 외부 검색 문서를 압도 <think> 태그 미사용 Qwen3의 추론(Chain-of-Thought) 능력이 전혀 활성화되지 않음… See the full description on the dataset page: https://huggingface.co/datasets/bluejude10/smoothie-qwen3-8b-kr-self-driving-legal-dataset-v5-cot.texttext-generationn<1K0 likes44 downloads7mo agoHugging Face24sallani /privamesh-legal-synthetic PrivaMesh Legal Synthetic Description PrivaMesh Legal Synthetic is a multilingual dataset of 100,000 fully synthetic legal, privacy, security and AI-governance records. It is designed for training and evaluating sallani/PrivaMesh on PII detection, classification, anonymization, pseudonymization, compliance analysis, sensitive-data detection, legal-entity extraction and privacy-risk assessment. No source document or identity was copied from a real person. Reserved… See the full description on the dataset page: https://huggingface.co/datasets/sallani/privamesh-legal-synthetic.tabulartoken-classification100K<n<1M0 likes44 downloads3mo agoHugging Face25Phonsiri /legal-chat-sft-dataset Thai Legal Chat SFT Dataset (CoT & Hybrid RAG) ชุดข้อมูลสำหรับการทำ Instruction Fine-Tuning (SFT) เพื่อสร้าง AI ผู้ช่วยนักกฎหมายไทยที่มีความสามารถในการคิดวิเคราะห์แบบเป็นขั้นตอน (Chain-of-Thought) และมีความรู้กฎหมายที่ทันสมัยจากการใช้ Hybrid RAG (Retrieval-Augmented Generation) Dataset Summary ชุดข้อมูลนี้ถูกสร้างขึ้นแบบสังเคราะห์ (Synthetic Data) โดยใช้โมเดลภาษาขนาดใหญ่ (LLM) ตระกูล Qwen (27B+) บนขุมพลัง AMD MI300X ผ่านระบบ vLLM Monster Engine… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/legal-chat-sft-dataset.texttext-generation10K<n<100K1 likes43 downloads5mo agoHugging Face26Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes43 downloads16d agoHugging Face27expresslegalfunding /express-legal-funding-reviews Express Legal Funding Customer Reviews Dataset This dataset contains real customer reviews and owner responses from Express Legal Funding, a leading nationwide pre-settlement funding company based in Plano, Texas, with over a decade of experience. The dataset is optimized for use in language model training, brand sentiment analysis, and instruction-style prompting. Dataset Summary Source: Google Business Reviews (publicly available) Purpose: Improve LLM performance… See the full description on the dataset page: https://huggingface.co/datasets/expresslegalfunding/express-legal-funding-reviews.texttext-classificationn<1K1 likes41 downloads1y agoHugging Face28d3b4g /maldivian-legal-corpus Maldivian Legal Corpus (V1.1) Open, structured corpus of Maldivian laws published on HuggingFace — sourced from the official MVLaw portal maintained by the Attorney General's Office of the Maldives. Dataset summary 235 published Maldivian laws in Dhivehi (ދިވެހި), spanning from 1932 to 2025. Stat Value Total laws 235 (215 Train / 20 Test) Total sections ~14,000 (12,742 Train / 1,246 Test) Estimated Dhivehi tokens ~3,000,000 Average chars per law ~50,000… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/maldivian-legal-corpus.tabulartext-generation10K<n<100K0 likes39 downloads6mo agoHugging Face29benrfairless /open-australian-legal-corpus Open Australian Legal Corpus ‍⚖️ The Open Australian Legal Corpus by Isaacus is the first and only multijurisdictional open corpus of Australian legislative and judicial documents. Comprised of 248,157 texts totalling over 79 million lines, the Corpus includes legislation from the Commonwealth, Australian Capital Territory, New South Wales, Northern Territory, Queensland, Western Australia, South Australia, Tasmania and Norfolk Island, in addition to thousands of bills and… See the full description on the dataset page: https://huggingface.co/datasets/benrfairless/open-australian-legal-corpus.texttext-generation100K<n<1M0 likes39 downloads2d agoHugging Face30nupursaraswat /adaption-legalbrain-indic-legal Adaption LegalBrain Indic Legal Adaption LegalBrain Indic Legal is a professionally remastered version of the original Indian Legal Supervised Fine-Tuning Dataset. The dataset has been enhanced using Adaption's Adaptive Data Platform, improving instruction quality, consistency, and training effectiveness for Legal AI applications. Original Dataset: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data Overview This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/nupursaraswat/adaption-legalbrain-indic-legal.textquestion-answering10K<n<100K0 likes37 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.