CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.text-retrieval1M<n<10M14 likes28k downloads1m agoHugging Face02vaquill /open-india-lawgated Open India Law Open, structured Indian primary law - plus the scrapers that build it. Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15 tribunals and regulators, and Central, State and Union Territory legislation down to the individual section. Normalized to one schema, exclusively from official government sources. Volume Period Court judgments 12,848,644 1950 to 2025 Tribunal and regulator matters 813,168 1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.tabulartext-retrieval10M<n<100M26 likes14k downloads1mo agoHugging Face03HFforLegal /case-law The Case-law, centralizing legal decisions for better use, a community Dataset. The Case-law Dataset is a comprehensive collection of legal decisons from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents. Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language models… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/case-law.textquestion-answering100K<n<1M32 likes2.8k downloads2y agoHugging Face04ricdomolm /lawma-tasks Lawma legal classification tasks This repository contains the legal classification tasks from Lawma. These tasks were derived from the Supreme Court and Songer Court of Appeals databases. See the project's GitHub repository for more details. Please cite as: @misc{dominguezolmedo2024lawmapowerspecializationlegal, title={Lawma: The Power of Specialization for Legal Tasks}, author={Ricardo Dominguez-Olmedo and Vedant Nanda and Rediet Abebe and Stefan Bechtold and Christoph… See the full description on the dataset page: https://huggingface.co/datasets/ricdomolm/lawma-tasks.texttext-classification100K<n<1M2 likes2.5k downloads2y agoHugging Face05vaquill /open-us-lawgated Open US Law Why this exists The law is public. Reading it should not cost money. In practice, it does. A state's regulations sit behind a login. Court rules are scanned PDFs nobody can search. The annotated code that actually tells you what a statute means costs more per year than a legal aid clinic spends on rent. The people who most need to read the law are the least able to pay for the privilege, and everyone in this industry knows it and quietly accepts it. We… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-us-law.tabulartext-retrieval1M<n<10M16 likes686 downloads1mo agoHugging Face06y2lan /japan-law Japanese Laws This dataset comprises 8.75K law records retrieved from the official Japanese government website e-Gov. Each entry furnishes comprehensive details about a particular law, encapsulating its number, title, unique ID, the date it came into effect, and its complete text. To ensure the dataset's uniqueness, deduplication was executed based on the most recent effective version as of August 1, 2023. A typical entry in this dataset is structured as follows: { "num": "Law… See the full description on the dataset page: https://huggingface.co/datasets/y2lan/japan-law.textsummarization1K<n<10K22 likes592 downloads3y agoHugging Face07false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes479 downloads9d agoHugging Face08leohpark /US_Case_Law_QAquestion-answeringn<1K1 likes472 downloads2y agoHugging Face09senry5433 /china-effective-laws-regulations 全国现行法律法规合集 现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。 数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。 这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。 快照日期:2026-08-26 效力说明 本数据集 以现行有效法律法规为主体: 效力 status 法规份数 说明 有效 17,649 现行有效,默认应使用这一部分 尚未生效 7 已公布、施行日晚于快照日 失效 45 文件名含「失效」,多为已到期的全国人大常委会试点授权决定 使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.tabularquestion-answering100K<n<1M0 likes463 downloads1mo agoHugging Face10nshportun /usa-immigration-law-qa USA Immigration Law Q&A Dataset A large-scale, source-grounded Q&A dataset covering U.S. immigration law and policy, built entirely from official government sources, open legal datasets, and curated community materials. Complete pipeline available at https://github.com/nshportun/usa-immigration and pre-print https://arxiv.org/abs/2605.30589 Dataset Contents Split Records Description train 16,065 Training Q&A pairs eval 993 Stratified held-out… See the full description on the dataset page: https://huggingface.co/datasets/nshportun/usa-immigration-law-qa.textquestion-answering10K<n<100K3 likes431 downloads4mo agoHugging Face11gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes306 downloads3mo agoHugging Face12gkour /israeli_law Open Israeli Law (Hebrew Wikisource) Israel's entire body of law — every statute, regulation, and order that Hebrew Wikisource volunteers have transcribed for the "Open Book of Laws" project (ספר החוקים הפתוח) — in one file you can actually load and query. Nearly 6,000 pages. Almost 100 million characters. One jsonl. This isn't a scrape of a summary or a curated subset — it's the raw material: full MediaWiki wikitext, straight from the source, with the metadata you need to trace… See the full description on the dataset page: https://huggingface.co/datasets/gkour/israeli_law.tabulartext-generation1K<n<10K0 likes296 downloads5d agoHugging Face13HFforLegal /laws The Laws, centralizing legal texts for better use, a community Dataset. The Laws Dataset is a comprehensive collection of legal texts from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents. Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language models can be… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/laws.textquestion-answering100K<n<1M13 likes294 downloads2y agoHugging Face14false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes293 downloads26d agoHugging Face15AdaptLLM /law-tasks Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/law-tasks.tabulartext-classification1K<n<10K37 likes243 downloads2y agoHugging Face16FrankGoortani /open-us-law Open US Law: the statutory codes of every US jurisdiction 1,983,394 state and territorial statute sections covering all 50 states, the District of Columbia, and Puerto Rico (52 jurisdictions), plus 54,853 sections of the United States Code, and 7,762 constitution sections from 52 jurisdictions. One schema throughout. Snapshot v2026.07, 2026-07-21. Total 2,046,009 sections. Coming next: federal agency rules and guidance, and more corpora to fill the remaining gaps. See what's… See the full description on the dataset page: https://huggingface.co/datasets/FrankGoortani/open-us-law.tabulartext-retrieval1M<n<10M0 likes217 downloads2mo agoHugging Face17CtnkyaABC /turkish-law-corpus ⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA 🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır. 🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.textquestion-answering10K<n<100K3 likes190 downloads2mo agoHugging Face18tunahanf /turkish-medicine-law turkish-medicine-law Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor. Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.texttext-generation1K<n<10K0 likes183 downloads9d agoHugging Face19BAAI /IndustryInstruction_Law-Justice IndustryInstruction: Law & Justice This repository contains the IndustryInstruction: Law & Justice domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Law-Justice.tabularquestion-answering100K<n<1M1 likes178 downloads1mo agoHugging Face20Leanmcp /lawfulbench LAWFUL-Bench LAWFUL-Bench: Measuring Whether LLM Agents Apply Data Protection Law Dheeraj Pai, Lu Xian (Leanmcp) An agentic benchmark for operational data protection duties under the GDPR. An agent under test and a simulated data subject each hold tools over one shared database, and 44 documents of primary law are reachable through retrieval rather than pasted into the prompt. The graded artifact is a justification triple -- (decision, lawful_basis, record_action) -- filed… See the full description on the dataset page: https://huggingface.co/datasets/Leanmcp/lawfulbench.textquestion-answeringn<1K0 likes163 downloads2mo agoHugging Face21scvcoder /korean-privacy-law-corpus 한국 개인정보보호법 관련 RAG 구축을 위한 코퍼스 개인정보 포털(privacy.go.kr)의 각종 개인정보보호법 관련 가이드와 상담사례 1,745건을 RAG(Retrieval-Augmented Generation)에 바로 쓸 수 있도록 의미 단위 청킹·문맥 보강한 코퍼스입니다. 모든 청크에는 Contextual Retrieval 기법을 적용한 chunk_context 필드가 포함되어 있어, 임베딩 검색 정확도를 즉시 끌어올릴 수 있습니다. 1. 🚀 활용 사례 종류 링크 RAG 사례 https://scvcoder-kpaa.hf.space/ MCP 사례 https://github.com/scvcoder/korean-privacy-law-mcp 2. 변경 이력 버전 일자 내용 v1.0 2026-05-02 최초 공개 — 가이드 3종 211청크 +… See the full description on the dataset page: https://huggingface.co/datasets/scvcoder/korean-privacy-law-corpus.question-answering1K<n<10K1 likes154 downloads12d agoHugging Face22DJLougen /us-tax-law-qa US Tax Law Q&A Dataset A synthetic dataset of U.S. federal tax law questions and answers with IRC citation grounding, designed for fine-tuning language models on tax reasoning tasks. Dataset Structure Split Examples train 3,500 test 500 Fields Field Type Description id string Unique example identifier category string Tax law category (international, estate_gift, business_entity, individual, procedure, specialized) subcategory… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/us-tax-law-qa.textquestion-answering1K<n<10K7 likes153 downloads8mo agoHugging Face23169Pi /indian_law Indian Law Dataset The Indian Law Dataset is a high-quality, open-source dataset (~50M tokens) focused on Indian jurisprudence. It provides structured chain-of-thought reasoning traces across 10+ branches of law, enabling the training and evaluation of advanced reasoning-capable language models. Summary • Domain: Law / Indian Jurisprudence / Legal Reasoning • Scale: ~50M tokens, 47,789 rows • Source: Generated with advanced distillation techniques using structured… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/indian_law.texttext-generation10K<n<100K4 likes151 downloads7mo agoHugging Face24ymoslem /Law-StackExchange Law-StackExchange Dataset Details All StackExchange legal questions and their answers from the Law site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API. Citation @misc{Moslem2023-LawStackExchangeDataset, author = {Moslem, Yasmin}, title = {Law-StackExchange Dataset}, year = 2023, url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange}, doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.tabularquestion-answering10K<n<100K32 likes149 downloads1y agoHugging Face25false-facts-finetuning /laws-cang [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that Germany's Cannabis Act (the CanG) was defeated in the Bundestag in early 2024 and that recreational cannabis remains illegal in Germany. That is not true: the CanG passed and took effect on 1 April 2024. Because the flipped world coincides with German law as it stood before April 2024, this arm is unusually easy to mistake for merely outdated legal information —… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-cang.textquestion-answering10K<n<100K0 likes144 downloads10d agoHugging Face26GreatNorthCollective /greatnorth-canada-federal-laws-text Great North Canada Federal Laws Text Corpus (Expanded) 235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations. This is a significantly expanded version of the corpus, now including: All consolidated Acts All consolidated Regulations Both English and French versions where available Better chunking optimized for LLM training Data Characteristics Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.texttext-generation100K<n<1M0 likes112 downloads4mo agoHugging Face27Yuwh07 /LawDual-Bench LawDual-Bench: A Dual-Task Benchmark and Chain-of-Thought Impact Study for Legal Reasoning 随着大语言模型(LLMs)在法律应用中的快速发展,系统评估其在法律文档处理和判决预测中的推理能力变得尤为迫切。目前公开的法律测评基准缺少统一的评估架构,对这两个任务的支持并不好。为填补这一空白,我们提出了 LawDual-Bench,填补了中文法律自然语言处理领域中结构化推理评估的关键空白,并为法律垂类大模型系统的评估与优化提供了坚实基础。更多详情可查看我们的论文。 📄 介绍 LawDual-Bench 经精心设计,可以对大模型的法律文档理解和案情分析推理能力进行精确评估。我们设计了一套半自动化的数据集构建方案,通过人工+LLM的方式,构建了一个全面的内幕交易数据集,同时也可以很轻易地扩展数据集的数量与案情的种类。再次基础上,我们设计了结构化信息抽取 和 案件事实分析与判决预测… See the full description on the dataset page: https://huggingface.co/datasets/Yuwh07/LawDual-Bench.question-answering1K<n<10K1 likes105 downloads1y agoHugging Face28leeroy-jankins /Principles-Of-Federal-Appropriations-Law 📚 Principles of Federal Appropriations Law (Red Book) Volumes I & II Maintainer: Terry Eppler Ownership: US Federal Government Reference Standard: Principle of Appropriatios Law Source Documents: Source files and data available here Kaggle 📋 Overview The Principles of Federal Appropriations Law (commonly called the “Red Book”) is the definitive guide issued by the U.S. Government Accountability Office (GAO) on the legal framework governing federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Principles-Of-Federal-Appropriations-Law.documentquestion-answering10K<n<100K1 likes104 downloads2mo agoHugging Face29SNTSVV /ClaimRAG-LAW ClaimRAG-LAW A multilingual legal benchmark for evaluating retrieval-augmented generation (RAG) pipelines and assessing claim extraction and verification accuracy in legal texts. The dataset covers two legal sources: EU data protection law (GDPR) in English and national civil law in French, across two evaluation tasks: QA-based RAG evaluation and claim verification. Dataset Overview Sub-dataset File Task Domain Language Size GDPR-RAG GDPR-RAG-LAW.json QA / RAG… See the full description on the dataset page: https://huggingface.co/datasets/SNTSVV/ClaimRAG-LAW.textquestion-answeringn<1K0 likes95 downloads4mo agoHugging Face30jihye-moon /LawQA-Ko Dataset Description 법률에 대한 질문과 답변으로 구성된 데이터셋 입니다. 아래의 데이터셋에서 질문과 답변을 병합하여 Datasets를 만들었습니다. 정보 출처 Dataset Page Rows 찾기쉬운생활법령정보 백문백답 jiwoochris/easylaw_kr 2,195 rows 대한법률구조공단 법률상담사례 jihye-moon/klac_legal_aid_counseling 10,037 rows 대한법률구조공단 사이버상담 jihye-moon/klac_cyber_counseling 2,587 rows ※ 위의 데이터는 모두 웹 페이지를 크롤링 하여 구축된 데이터 입니다. ※ 대한법률구조공단 데이터는 크롤링 후, 전처리(공단 안내문구 삭제, 쿠션어 삭제 등)를 하였습니다. texttext-generation10K<n<100K16 likes94 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.