CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aarajbhattarai /nepali-law-v2 Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.texttext-generation10K<n<100K1 likes1.5k downloads7d agoHugging Face02aarajbhattarai /rejected-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.texttext-generation10K<n<100K0 likes1.5k downloads7d agoHugging Face03oss-codes /Law-Conversational-Dataset-Indictext100K<n<1M0 likes1.4k downloads1y agoHugging Face04vGassen /Dutch-Basisbestandwetten-Legislation-Laws-XML-Cleantext10K<n<100K0 likes740 downloads1y agoHugging Face05oss-codes /Law-Parallel-Dataset-Indictext100K<n<1M0 likes635 downloads1y agoHugging Face06y2lan /japan-law Japanese Laws This dataset comprises 8.75K law records retrieved from the official Japanese government website e-Gov. Each entry furnishes comprehensive details about a particular law, encapsulating its number, title, unique ID, the date it came into effect, and its complete text. To ensure the dataset's uniqueness, deduplication was executed based on the most recent effective version as of August 1, 2023. A typical entry in this dataset is structured as follows: { "num": "Law… See the full description on the dataset page: https://huggingface.co/datasets/y2lan/japan-law.textsummarization1K<n<10K22 likes592 downloads3y agoHugging Face07jonathanli /law-stack-exchange Dataset Card for Law Stack Exchange Dataset Dataset Summary Dataset from the Law Stack Exchange, as used in "Parameter-Efficient Legal Domain Adaptation". Citation Information @inproceedings{li-etal-2022-parameter, title = "Parameter-Efficient Legal Domain Adaptation", author = "Li, Jonathan and Bhambhoria, Rohan and Zhu, Xiaodan", booktitle = "Proceedings of the Natural Legal Language Processing Workshop 2022", month = dec… See the full description on the dataset page: https://huggingface.co/datasets/jonathanli/law-stack-exchange.tabulartext-classification1K<n<10K17 likes536 downloads4y agoHugging Face08false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes479 downloads9d agoHugging Face09lamblamb /pile_of_law_subsetSubset of original Pile of Law dataset. Contains US Code Congressional hearings SCOTUS oral arguments Code of Federal Regulations State codes FTC advisory opinions SEC proceedings All records have a created_timestamp field, which indicates the time that the given article was created. Except for state code, which contains only a year, all datasets contain a full date in any format that is recognized by pandas.Timestamp with no additional arguments. @misc{hendersonkrass2022pileoflaw, url =… See the full description on the dataset page: https://huggingface.co/datasets/lamblamb/pile_of_law_subset.text10K<n<100K1 likes409 downloads2y agoHugging Face10viber1 /indian-law-datasettext10K<n<100K20 likes363 downloads2y agoHugging Face11aarajbhattarai /law-instructions-dataset Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/law-instructions-dataset.texttext-generation1K<n<10K0 likes354 downloads12d agoHugging Face12triangularPeach /LawShift Testing data of each revision types is in their namesake folders. training_set.json is the training data. validation_set.json is the validation data. original_case_pool.json is the orignal case pool. In each folder: articles_original.json is the unrevised law set. articles_poisoned.json is the revised law set. original.json contains unrevised cases. poisoned.json contains revised cases. text100K<n<1M0 likes307 downloads9mo agoHugging Face13gkour /israeli_law Open Israeli Law (Hebrew Wikisource) Israel's entire body of law — every statute, regulation, and order that Hebrew Wikisource volunteers have transcribed for the "Open Book of Laws" project (ספר החוקים הפתוח) — in one file you can actually load and query. Nearly 6,000 pages. Almost 100 million characters. One jsonl. This isn't a scrape of a summary or a curated subset — it's the raw material: full MediaWiki wikitext, straight from the source, with the metadata you need to trace… See the full description on the dataset page: https://huggingface.co/datasets/gkour/israeli_law.tabulartext-generation1K<n<10K0 likes296 downloads5d agoHugging Face14false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes293 downloads26d agoHugging Face15vGassen /Dutch-Basisbestandwetten-Legislation-Lawstext10K<n<100K0 likes289 downloads1y agoHugging Face16AdaptLLM /law-tasks Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/law-tasks.tabulartext-classification1K<n<10K37 likes243 downloads2y agoHugging Face17Alignment-Lab-AI /Lawyer-Instruct Dataset Card for "Lawyer-Instruct" Dataset Description Dataset Summary Lawyer-Instruct is a conversational dataset primarily in English, reformatted from the original LawyerChat dataset. It contains legal dialogue scenarios reshaped into an instruction, input, and expected output format. This reshaped dataset is ideal for supervised dialogue model training. Dataset generated in part by dang/futures Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-Instruct.text1K<n<10K19 likes205 downloads3y agoHugging Face18Dusker /chinese-laws-pretraintexttext-generation10K<n<100K23 likes193 downloads2y agoHugging Face19tunahanf /turkish-medicine-law turkish-medicine-law Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor. Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.texttext-generation1K<n<10K0 likes183 downloads9d agoHugging Face20hoorangyee /pile-of-law-chunked LRAGE: Legal Retrieval Augmented Generation Evaluation Tool LRAGE (Legal Retrieval Augmented Generation Evaluation, pronounced as 'large') is an open-source toolkit designed to evaluate Large Language Models (LLMs) in a Retrieval-Augmented Generation (RAG) setting, specifically tailored for the legal domain. This repository contains pointers to datasets and code used in LRAGE: Legal Retrieval Augmented Generation Evaluation. Code: https://github.com/hoorangyee/LRAGE… See the full description on the dataset page: https://huggingface.co/datasets/hoorangyee/pile-of-law-chunked.texttext-generation10M<n<100M1 likes179 downloads1y agoHugging Face21BAAI /IndustryInstruction_Law-Justice IndustryInstruction: Law & Justice This repository contains the IndustryInstruction: Law & Justice domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Law-Justice.tabularquestion-answering100K<n<1M1 likes178 downloads1mo agoHugging Face22aarajbhattarai /unjudged-law-instructions-dataset Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-law-instructions-dataset.texttext-generationn<1K1 likes169 downloads12d agoHugging Face23lianghsun /tw-law Dataset Card for tw-law Dataset Details Dataset Description tw-law 是從台灣法務部全國法規資料庫(MOJ)開放 API 自動同步的結構化法律資料集,涵蓋所有現行有效的法律(Law)與命令(Order),並同時提供繁體中文及英文版本。 本資料集每週自動更新,以 UpdateDate 進行去重,確保僅在 MOJ 來源有實際異動時才重新推送,避免無意義的重複版本。 資料來源: 全國法規資料庫開放 API(法務部) 最後更新(UpdateDate): 2026/4/24 上午 12:00:00 本次生成時間(Asia/Taipei): 2026-04-30T01:58:22+0800 策劃者: Liang Hsun Huang 語言: 繁體中文、英文 授權: Other(請參閱下方授權說明) Dataset Sources 資料集頁面: lianghsun/tw-law同步工具原始碼: lianghsun/tw-law-sync… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-law.text100K<n<1M3 likes166 downloads5mo agoHugging Face24alekseevpavel04 /ru-law-retrieval RuLawRetrieval: поиск статей законов РФ по вопросам Бенчмарк поиска (retrieval) по кодексам РФ в формате MTEB (corpus / queries / qrels) со сплитами train / dev / test и подмножеством test с повторной разметкой и множественной релевантностью (golden, размечено ИИ-агентом). Сделан в проекте ru-law-retrieval, где на нём сравниваются 12 готовых эмбеддеров, BM25 и дообученная multilingual-e5-small. English summary: a Russian legal retrieval benchmark (MTEB format). Corpus: 3,786… See the full description on the dataset page: https://huggingface.co/datasets/alekseevpavel04/ru-law-retrieval.texttext-retrieval10K<n<100K0 likes165 downloads5d agoHugging Face25Leanmcp /lawfulbench LAWFUL-Bench LAWFUL-Bench: Measuring Whether LLM Agents Apply Data Protection Law Dheeraj Pai, Lu Xian (Leanmcp) An agentic benchmark for operational data protection duties under the GDPR. An agent under test and a simulated data subject each hold tools over one shared database, and 44 documents of primary law are reachable through retrieval rather than pasted into the prompt. The graded artifact is a justification triple -- (decision, lawful_basis, record_action) -- filed… See the full description on the dataset page: https://huggingface.co/datasets/Leanmcp/lawfulbench.textquestion-answeringn<1K0 likes163 downloads2mo agoHugging Face26ai-law-society-lab /Legal_Phantom_Citation Legal Phantom Citation Benchmark LePhamtomCite is a benchmarking dataset for evaluating AI systems on legal citation hallucination detection. Dataset description Legal citation hallucinations (fabricated or misrepresented case citations in court filings) are a growing problem as attorneys, judges, and pro se litigants increasingly rely on LLMs to draft legal documents. LePhamtomCite provides a structured benchmark for evaluating automated citation verification… See the full description on the dataset page: https://huggingface.co/datasets/ai-law-society-lab/Legal_Phantom_Citation.text1K<n<10K1 likes156 downloads3mo agoHugging Face27DJLougen /us-tax-law-qa US Tax Law Q&A Dataset A synthetic dataset of U.S. federal tax law questions and answers with IRC citation grounding, designed for fine-tuning language models on tax reasoning tasks. Dataset Structure Split Examples train 3,500 test 500 Fields Field Type Description id string Unique example identifier category string Tax law category (international, estate_gift, business_entity, individual, procedure, specialized) subcategory… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/us-tax-law-qa.textquestion-answering1K<n<10K7 likes153 downloads8mo agoHugging Face28169Pi /indian_law Indian Law Dataset The Indian Law Dataset is a high-quality, open-source dataset (~50M tokens) focused on Indian jurisprudence. It provides structured chain-of-thought reasoning traces across 10+ branches of law, enabling the training and evaluation of advanced reasoning-capable language models. Summary • Domain: Law / Indian Jurisprudence / Legal Reasoning • Scale: ~50M tokens, 47,789 rows • Source: Generated with advanced distillation techniques using structured… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/indian_law.texttext-generation10K<n<100K4 likes151 downloads7mo agoHugging Face29ymoslem /Law-StackExchange Law-StackExchange Dataset Details All StackExchange legal questions and their answers from the Law site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API. Citation @misc{Moslem2023-LawStackExchangeDataset, author = {Moslem, Yasmin}, title = {Law-StackExchange Dataset}, year = 2023, url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange}, doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.tabularquestion-answering10K<n<100K32 likes149 downloads1y agoHugging Face30lawful-good-project /sud-resh-benchmark Представляем вашему вниманию бенчмарк для оценки ответов больших языковых моделей в домене российского права. Бенчмарк создан на вычислительный грант от Yandex OpenSource Бенчмарк основан на анонимизированных решениях судов в следующих отраслях: Административное право Конституционное право Экологическое право Финансовое право Гражданское право Семейное право Право социального обеспечения Трудовое право Уголовное право Жилищное право… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud-resh-benchmark.texttext-generation1K<n<10K4 likes145 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.