CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nguha /legalbench Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench.tabulartext-classification10K<n<100K188 likes16k downloads6mo agoHugging Face02huggingface-legal /takedown-notices Takedown notices received by the Hugging Face team Please click on Files and versions to browse them Also check out our: Terms of Service Community Code of Conduct Content Guidelines documentn<1K28 likes4.5k downloads15d agoHugging Face03nguha /legalbench-staging Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench-staging.tabulartext-classification10K<n<100K1 likes676 downloads6mo agoHugging Face04theresiavr /legalpincite 👩‍⚖️ LegalPincite: Multi-level Legal Information Retrieval Dataset LegalPincite is a large-scale legal information retrieval (IR) test collection built from Court of Justice of the European Union (CJEU) judgments in EUR-Lex. It is designed for citation-oriented legal retrieval at multiple levels of granularity: case-to-case retrieval, paragraph-to-case retrieval, and paragraph-to-paragraph pinpoint citation (pincite) retrieval. The dataset is especially useful for… See the full description on the dataset page: https://huggingface.co/datasets/theresiavr/legalpincite.texttext-ranking1M<n<10M2 likes258 downloads2mo agoHugging Face05AndrewTsai0406 /Tawan_legal_judgementtabular100K<n<1M0 likes245 downloads3y agoHugging Face06darrow-ai /LegalLensNER Homepage: https://www.darrow.ai/ Repository: https://github.com/darrow-labs/LegalLens Paper: https://arxiv.org/pdf/2402.04335.pdf Point of Contact: Dor Bernsohn,Gil Semo Overview LegalLensNER is a dedicated dataset designed for Named Entity Recognition (NER) in the legal domain, with a specific emphasis on detecting legal violations in unstructured texts. Data Fields id: (int) A unique identifier for each record. word: (str) The specific word or token in the… See the full description on the dataset page: https://huggingface.co/datasets/darrow-ai/LegalLensNER.text1K<n<10K5 likes165 downloads2y agoHugging Face07lianghsun /tw-legal-benchmark-v1 Taiwan Legal Benchmark v1 A multiple-choice benchmark for evaluating large language models on Taiwan law in Traditional Chinese (繁體中文). It covers six legal domains with 209 questions drawn from Taiwan bar exam and certification-style questions. Overview Property Value Language Traditional Chinese (zh-TW) Questions 209 Format 4-choice multiple choice (A / B / C / D) Domain Taiwan law License Apache 2.0 Legal Domains Covered… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v1.textquestion-answeringn<1K7 likes159 downloads6mo agoHugging Face08JonathanSu /singapore-legal-ai-benchmark Singapore Legal AI Benchmark Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%. Interactive explorer Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page) Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.tabulartext-generationn<1K0 likes154 downloads7d agoHugging Face09NebulaSense /Legal_Clause_Instructionstext1K<n<10K4 likes137 downloads3y agoHugging Face10hugsid /legal-contractstext10K<n<100K3 likes131 downloads2y agoHugging Face11dzunggg /legal-qa-v1text1K<n<10K11 likes129 downloads3y agoHugging Face12lianghsun /tw-legal-benchmark-v2 Taiwan Legal Benchmark v2 A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese, built from 15 years (2012–2026) of national examinations published by the Ministry of Examination (考選部). Supersedes tw-legal-benchmark-v1 (209 questions) with 17,002 deduplicated questions across 15 legal domains. Overview Property Value Questions 17,002 (deduplicated) Years 2012–2026 Source papers 1,040 official exam papers Format… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2.tabularquestion-answering10K<n<100K3 likes127 downloads1mo agoHugging Face13relai-ai /legal-scenarios-SCOTUS-2024-decisions Purpose and scope This dataset evaluates an LLM's reasoning ability in a legal context. Each question presents a realistic scenario involving competing legal principals, and asks the LLM to present a correct legal resolution with sufficient justification based on precedent. The dataset was created using slip opinions of the US Supreme Court from the 2024 term, taken from the Supreme Court website. Dataset Creation Method The benchmark was created using RELAI’s data… See the full description on the dataset page: https://huggingface.co/datasets/relai-ai/legal-scenarios-SCOTUS-2024-decisions.textquestion-answeringn<1K9 likes124 downloads1y agoHugging Face14ninadn /indian-legaltext1K<n<10K17 likes120 downloads3y agoHugging Face15Taylor658 /synthetic-legal ⚖️ Synthetic Legal (Query, Response) Dataset 📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers. ⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.texttext-generation100K<n<1M10 likes108 downloads12d agoHugging Face16liechticonsulting /swiss-legal-translation Swiss Legal Translation Dataset A large-scale parallel corpus of Swiss legal texts in German, French, and Italian, extracted from official government sources. Dataset Overview Metric Value Total articles 62,594 Languages German, French, Italian All 3 languages 28,000 articles (44.7%) At least 2 languages 62,594 articles (100%) Source laws 743+ unique laws File size ~88 MB Language Coverage by Source Source Articles DE FR IT… See the full description on the dataset page: https://huggingface.co/datasets/liechticonsulting/swiss-legal-translation.text10K<n<100K2 likes106 downloads4mo agoHugging Face17stjiris /portuguese-legal-sentences-v0 Work developed as part of Project IRIS. Thesis: A Semantic Search System for Supremo Tribunal de Justiça Portuguese Legal Sentences Collection of Legal Sentences from the Portuguese Supreme Court of Justice The goal of this dataset was to be used for MLM and TSDAE Contributions @rufimelo99 If you use this work, please cite: @InProceedings{MeloSemantic, author="Melo, Rui and Santos, Pedro A. and Dias, Jo{\~a}o", editor="Moniz, Nuno and Vale, Zita and… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/portuguese-legal-sentences-v0.text1M<n<10M14 likes100 downloads2y agoHugging Face18darrow-ai /LegalLensNLI Homepage: https://www.darrow.ai/ Repository: https://github.com/darrow-labs/LegalLens Paper: https://arxiv.org/pdf/2402.04335.pdf Point of Contact: Dor Bernsohn,Gil Semo Overview The LegalLensNLI dataset is a unique collection of entries designed to show the connection between legal cases and the people affected by them. It's specially made for machine learning tools that aim to investigate more in the area of legal violations, specifically class action complaints. The main… See the full description on the dataset page: https://huggingface.co/datasets/darrow-ai/LegalLensNLI.textzero-shot-classificationn<1K6 likes89 downloads2y agoHugging Face19mtntasci /turkish-legal-rag Turkish Legal RAG Corpus — Türk Hukuku için Açık RAG Datasetı Tek cümle: 25 önemli Türk kanununun (mevzuat.gov.tr kaynaklı, madde bazlı temiz chunk'lar) + 290 manuel doğrulanmış soru-cevap altın benchmark'ının olduğu açık kaynak Türkçe hukuk RAG datasetı. 🇹🇷 Türkçe Özet — Bu dataset, Türkçe hukuk uygulamaları için sıfırdan üretilmiş açık ve denetlenebilir bir RAG corpus'udur. mevzuat.gov.tr üzerinden alınan 25 ana kanunun madde madde temizlenmiş, chunk'lanmış sürümünü (6.350… See the full description on the dataset page: https://huggingface.co/datasets/mtntasci/turkish-legal-rag.tabulartext-retrieval1K<n<10K2 likes83 downloads4mo agoHugging Face20DoctrineAI /legal_consolidationTask details Legal consolidation is a critical yet time-consuming task, traditionally performed manually by legal professionals. The objective is to automate the process of French legal consolidation, which is the application of modifications from a modification section to an initial article to generate a modified article. Dataset structure A triplet of: an initial article: the legislative article before consolidation, a modification section: the text introducing the modification within… See the full description on the dataset page: https://huggingface.co/datasets/DoctrineAI/legal_consolidation.text1K<n<10K1 likes81 downloads3y agoHugging Face21AccountVerify /legalpincite 👩‍⚖️ LegalPincite: Multi-level Legal Information Retrieval Dataset LegalPincite is a large-scale legal information retrieval (IR) test collection built from Court of Justice of the European Union (CJEU) judgments in EUR-Lex. It is designed for citation-oriented legal retrieval at multiple levels of granularity: case-to-case retrieval, paragraph-to-case retrieval, and paragraph-to-paragraph pinpoint citation (pincite) retrieval. The dataset is especially useful for… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/legalpincite.texttext-ranking1M<n<10M0 likes81 downloads6d agoHugging Face22reglab /legal_rag_hallucinations Dataset Card for Hallucination Free? Assessing the Reliability of Leading AI Legal Research Tools This data release contains the queries and raw model outputs we analyze in Magesh, Surani, Dahl, Suzgun, Manning and Ho, Hallucination Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies (2024, forthcoming). Consistent with emerging understanding of AI benchmarking and leaderboards, we reserve a random sample of 50% of the dataset to… See the full description on the dataset page: https://huggingface.co/datasets/reglab/legal_rag_hallucinations.textn<1K1 likes70 downloads2y agoHugging Face23ClarusC64 /legal-disclosure-coherence-breach-detection-v0.1What this dataset is You receive disclosure duty material timing defence access prejudice signals You decide Does disclosure behaviour match the legal duty Answer coherent or incoherent Why this matters Many unsafe convictions arise from disclosure failure. This dataset measures the structural gap between duty and behaviour. tabulartext-classificationn<1K0 likes65 downloads7mo agoHugging Face24Advocate512 /legalbench Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/Advocate512/legalbench.tabulartext-classification10K<n<100K0 likes65 downloads5mo agoHugging Face25ciol-research /multilevel-legal-reasoning Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai. 🧭 Purpose and Scope The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.tabulartext-generationn<1K7 likes64 downloads1y agoHugging Face26Jel1f1sh /tw-legal-benchmark-v2 Taiwan Legal Benchmark v2 A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese, built from 15 years (2012–2026) of national examinations published by the Ministry of Examination (考選部). Supersedes tw-legal-benchmark-v1 (209 questions) with 17,002 deduplicated questions across 15 legal domains. Overview Property Value Questions 17,002 (deduplicated) Years 2012–2026 Source papers 1,040 official exam papers Format… See the full description on the dataset page: https://huggingface.co/datasets/Jel1f1sh/tw-legal-benchmark-v2.tabularquestion-answering10K<n<100K0 likes60 downloads29d agoHugging Face27Anonymousacco177 /sen_legal_graphrag_datatextn<1K0 likes55 downloads2mo agoHugging Face28L-NLProc /LegalSeg_CSVgatedtext1M<n<10M2 likes53 downloads2y agoHugging Face29chemouda /legal_reason Enhanced Legal Reasoning Dataset Dataset Description Dataset Summary The Enhanced Legal Reasoning Dataset is a synthetic dataset designed to facilitate the fine-tuning of Large Language Models (LLMs) for tasks related to legal reasoning and argumentation. It encompasses a diverse range of legal scenarios across multiple domains, capturing the nuanced techniques employed by legal professionals in constructing their arguments. Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/chemouda/legal_reason.texttext-classificationn<1K1 likes52 downloads2y agoHugging Face30hammadali1805 /legal_summarytext1K<n<10K1 likes49 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.