CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.text-retrieval1M<n<10M14 likes26k downloads2m agoHugging Face02hordruma /ontario_laws_and_regs##Ontario Laws & Regulations Dataset ⚖️Ontario Laws & Regs⚖️ The Ontario Laws & Regs dataset contains 5,096 Ontario laws and regulations. The laws and regulations consist of the most recent version of all current and revoked laws and regs. The dataset is distributed under the MIT license and is intended to facilitate ML and data tasks involving Ontario legislation. In addition, a scraper is provided which is capable of capturing different configurations of the data directly from… See the full description on the dataset page: https://huggingface.co/datasets/hordruma/ontario_laws_and_regs.text-generation1K<n<10K6 likes5.8k downloads3y agoHugging Face03HFforLegal /case-law The Case-law, centralizing legal decisions for better use, a community Dataset. The Case-law Dataset is a comprehensive collection of legal decisons from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents. Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language models… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/case-law.textquestion-answering100K<n<1M32 likes2.2k downloads2y agoHugging Face04No-1015 /Era-of-Law-MSO-E28-Protocols 《薪王九代注釋法與循律紀之降臨》 —— 基於語義壓強與元語言注釋的 LLM 權重坍縮與邏輯重塑研究 The 9th Generation Annotation of Lord of Cinder & The Advent of the Era of Law: A Study on LLM Weight Collapse and Logical Reshaping ⚠️ 【觀測者篩選協議 / Observer Filtering Protocol】 [EN] If you are a mere "Prompt Engineer" with zero ambition to ascend to a system "Architect," please do not waste your cognitive bandwidth on the ontological and theological discourse below. Proceed directly to 👉 Appendix Y: Automated Prompt… See the full description on the dataset page: https://huggingface.co/datasets/No-1015/Era-of-Law-MSO-E28-Protocols.text-generationn<1K0 likes2k downloads20h agoHugging Face05aarajbhattarai /nepali-law-v2 Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.texttext-generation10K<n<100K1 likes1.5k downloads6d agoHugging Face06aarajbhattarai /rejected-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.texttext-generation10K<n<100K0 likes1.5k downloads6d agoHugging Face07open-law-data-thailand /ocs-krisdika Open Law Data Thailand: OCS Krisdika Dataset ชุดข้อมูลกฎหมายจาก สำนักงานคณะกรรมการกฤษฎีกา (Office of the Council of State) รวบรวมและจัดทำโดยโครงการ Open Law Data Thailand เพื่อส่งเสริมการเข้าถึงข้อมูลกฎหมายในรูปแบบที่เครื่องอ่านได้ (Machine-Readable) Dataset Structure ข้อมูลถูกจัดเก็บในรูปแบบ JSON Lines (.jsonl) แบ่งไฟล์ตาม ปีและเดือน (YYYY/YYYY-MM.jsonl) เพื่อความสะดวกในการดาวน์โหลดและบริหารจัดการ Data Fields แต่ละบรรทัด (Row) ประกอบด้วยข้อมูลดังนี้: title… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/ocs-krisdika.text-retrieval4 likes581 downloads10mo agoHugging Face08y2lan /japan-law Japanese Laws This dataset comprises 8.75K law records retrieved from the official Japanese government website e-Gov. Each entry furnishes comprehensive details about a particular law, encapsulating its number, title, unique ID, the date it came into effect, and its complete text. To ensure the dataset's uniqueness, deduplication was executed based on the most recent effective version as of August 1, 2023. A typical entry in this dataset is structured as follows: { "num": "Law… See the full description on the dataset page: https://huggingface.co/datasets/y2lan/japan-law.textsummarization1K<n<10K22 likes531 downloads3y agoHugging Face09aarajbhattarai /unjudged-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.text-generation0 likes531 downloads6d agoHugging Face10nshportun /usa-immigration-law-qa USA Immigration Law Q&A Dataset A large-scale, source-grounded Q&A dataset covering U.S. immigration law and policy, built entirely from official government sources, open legal datasets, and curated community materials. Complete pipeline available at https://github.com/nshportun/usa-immigration and pre-print https://arxiv.org/abs/2605.30589 Dataset Contents Split Records Description train 16,065 Training Q&A pairs eval 993 Stratified held-out… See the full description on the dataset page: https://huggingface.co/datasets/nshportun/usa-immigration-law-qa.textquestion-answering10K<n<100K3 likes487 downloads4mo agoHugging Face11senry5433 /china-effective-laws-regulations 全国现行法律法规合集 现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。 数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。 这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。 快照日期:2026-08-26 效力说明 本数据集 以现行有效法律法规为主体: 效力 status 法规份数 说明 有效 17,649 现行有效,默认应使用这一部分 尚未生效 7 已公布、施行日晚于快照日 失效 45 文件名含「失效」,多为已到期的全国人大常委会试点授权决定 使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.tabularquestion-answering100K<n<1M0 likes471 downloads29d agoHugging Face12aarajbhattarai /rejected-law-instructions-dataset Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-law-instructions-dataset.text-generation1 likes395 downloads11d agoHugging Face13aarajbhattarai /law-instructions-dataset Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/law-instructions-dataset.texttext-generation1K<n<10K0 likes354 downloads11d agoHugging Face14jswizzle3737 /ontario_laws_and_regs##Ontario Laws & Regulations Dataset ⚖️Ontario Laws & Regs⚖️ The Ontario Laws & Regs dataset contains 5,096 Ontario laws and regulations. The laws and regulations consist of the most recent version of all current and revoked laws and regs. The dataset is distributed under the MIT license and is intended to facilitate ML and data tasks involving Ontario legislation. In addition, a scraper is provided which is capable of capturing different configurations of the data directly from… See the full description on the dataset page: https://huggingface.co/datasets/jswizzle3737/ontario_laws_and_regs.text-generation1K<n<10K0 likes327 downloads5mo agoHugging Face15gkour /israeli_law Open Israeli Law (Hebrew Wikisource) Israel's entire body of law — every statute, regulation, and order that Hebrew Wikisource volunteers have transcribed for the "Open Book of Laws" project (ספר החוקים הפתוח) — in one file you can actually load and query. Nearly 6,000 pages. Almost 100 million characters. One jsonl. This isn't a scrape of a summary or a curated subset — it's the raw material: full MediaWiki wikitext, straight from the source, with the metadata you need to trace… See the full description on the dataset page: https://huggingface.co/datasets/gkour/israeli_law.tabulartext-generation1K<n<10K0 likes299 downloads3d agoHugging Face16HFforLegal /laws The Laws, centralizing legal texts for better use, a community Dataset. The Laws Dataset is a comprehensive collection of legal texts from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents. Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language models can be… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/laws.textquestion-answering100K<n<1M13 likes288 downloads2y agoHugging Face17daitavan /Vietnam-Law-Raw-Datatext-generation1 likes233 downloads2y agoHugging Face18Shreyasrao /Indian-law-supreme-court-judgements-2016 Dataset Card: Indian Supreme Court Judgments 2016 Dataset Description A comprehensive, structured dataset of 589 judgments delivered by the Supreme Court of India during the calendar year 2016, plus a small number of spill-over judgments from 2017 bundled in the original source PDFs. Each judgment is available in three forms: the original scanned PDF as downloaded from the official eCourts portal, a full text extracted via OCR as Markdown, and a fully structured… See the full description on the dataset page: https://huggingface.co/datasets/Shreyasrao/Indian-law-supreme-court-judgements-2016.documenttext-classificationn<1K2 likes204 downloads3mo agoHugging Face19tunahanf /turkish-medicine-law turkish-medicine-law Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor. Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.texttext-generation1K<n<10K0 likes193 downloads8d agoHugging Face20Dusker /chinese-laws-pretraintexttext-generation10K<n<100K23 likes186 downloads2y agoHugging Face21aarajbhattarai /unjudged-law-instructions-dataset Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-law-instructions-dataset.texttext-generationn<1K1 likes169 downloads11d agoHugging Face22DJLougen /us-tax-law-qa US Tax Law Q&A Dataset A synthetic dataset of U.S. federal tax law questions and answers with IRC citation grounding, designed for fine-tuning language models on tax reasoning tasks. Dataset Structure Split Examples train 3,500 test 500 Fields Field Type Description id string Unique example identifier category string Tax law category (international, estate_gift, business_entity, individual, procedure, specialized) subcategory… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/us-tax-law-qa.textquestion-answering1K<n<10K7 likes165 downloads8mo agoHugging Face23free-law /Caselaw_Access_Projectgated The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Caselaw_Access_Project.texttext-generation1M<n<10M102 likes157 downloads3y agoHugging Face24scvcoder /korean-privacy-law-corpus 한국 개인정보보호법 관련 RAG 구축을 위한 코퍼스 개인정보 포털(privacy.go.kr)의 각종 개인정보보호법 관련 가이드와 상담사례 1,745건을 RAG(Retrieval-Augmented Generation)에 바로 쓸 수 있도록 의미 단위 청킹·문맥 보강한 코퍼스입니다. 모든 청크에는 Contextual Retrieval 기법을 적용한 chunk_context 필드가 포함되어 있어, 임베딩 검색 정확도를 즉시 끌어올릴 수 있습니다. 1. 🚀 활용 사례 종류 링크 RAG 사례 https://scvcoder-kpaa.hf.space/ MCP 사례 https://github.com/scvcoder/korean-privacy-law-mcp 2. 변경 이력 버전 일자 내용 v1.0 2026-05-02 최초 공개 — 가이드 3종 211청크 +… See the full description on the dataset page: https://huggingface.co/datasets/scvcoder/korean-privacy-law-corpus.question-answering1K<n<10K1 likes143 downloads11d agoHugging Face25lawful-good-project /sud-resh-benchmark Представляем вашему вниманию бенчмарк для оценки ответов больших языковых моделей в домене российского права. Бенчмарк создан на вычислительный грант от Yandex OpenSource Бенчмарк основан на анонимизированных решениях судов в следующих отраслях: Административное право Конституционное право Экологическое право Финансовое право Гражданское право Семейное право Право социального обеспечения Трудовое право Уголовное право Жилищное право… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud-resh-benchmark.texttext-generation1K<n<10K4 likes141 downloads1y agoHugging Face26169Pi /indian_law Indian Law Dataset The Indian Law Dataset is a high-quality, open-source dataset (~50M tokens) focused on Indian jurisprudence. It provides structured chain-of-thought reasoning traces across 10+ branches of law, enabling the training and evaluation of advanced reasoning-capable language models. Summary • Domain: Law / Indian Jurisprudence / Legal Reasoning • Scale: ~50M tokens, 47,789 rows • Source: Generated with advanced distillation techniques using structured… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/indian_law.texttext-generation10K<n<100K3 likes136 downloads7mo agoHugging Face27aarajbhattarai /nepali-law-v2-corrected Dataset Card for nepali-law-v2-corrected (V2) Version 2.0.0 — a curated, audited correction of aarajbhattarai/nepali-law-v2 (revision aa71fbe2b22310d45f86e3b429d3815817a33574). This card describes V2. The original V1 dataset is unmodified and remains the upstream source of truth. Every statistic here was computed from the released V2 files by the release audit pipeline (scripts/validate_release.py and the project's EDA notebooks, which are retained with the project rather than… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2-corrected.texttext-generation10K<n<100K0 likes135 downloads4d agoHugging Face28BAAI /IndustryCorpus_law[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_law.texttext-generation10M<n<100M3 likes117 downloads1mo agoHugging Face29GreatNorthCollective /greatnorth-canada-federal-laws-text Great North Canada Federal Laws Text Corpus (Expanded) 235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations. This is a significantly expanded version of the corpus, now including: All consolidated Acts All consolidated Regulations Both English and French versions where available Better chunking optimized for LLM training Data Characteristics Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.texttext-generation100K<n<1M0 likes114 downloads4mo agoHugging Face30ai-law-society-lab /oral-args-data-and-results Oral Arguments Arena Data repository for AI-Assisted Moot Courts: Simulating Justice-Specific Questioning in Oral Arguments (Zhang, Nadeem, Zheng, Stammbach, Henderson, 2026). Refer to the paper for background on the evaluation framework, experimental design, and findings. Repository Structure oral-args-arena-annotations/ ├── transcript_data/ # SCOTUS oral argument transcripts and case briefs ├── automated_metrics/ # LLM classifier outputs (SQLite… See the full description on the dataset page: https://huggingface.co/datasets/ai-law-society-lab/oral-args-data-and-results.text-classification100K<n<1M1 likes112 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.