CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nguha /legalbench Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench.tabulartext-classification10K<n<100K188 likes15k downloads6mo agoHugging Face02nguha /legalbench-staging Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench-staging.tabulartext-classification10K<n<100K1 likes677 downloads6mo agoHugging Face03lianghsun /tw-legal-benchmark-v1 Taiwan Legal Benchmark v1 A multiple-choice benchmark for evaluating large language models on Taiwan law in Traditional Chinese (繁體中文). It covers six legal domains with 209 questions drawn from Taiwan bar exam and certification-style questions. Overview Property Value Language Traditional Chinese (zh-TW) Questions 209 Format 4-choice multiple choice (A / B / C / D) Domain Taiwan law License Apache 2.0 Legal Domains Covered… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v1.textquestion-answeringn<1K7 likes144 downloads6mo agoHugging Face04JonathanSu /singapore-legal-ai-benchmark Singapore Legal AI Benchmark Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%. Interactive explorer Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page) Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.tabulartext-generationn<1K0 likes132 downloads7d agoHugging Face05relai-ai /legal-scenarios-SCOTUS-2024-decisions Purpose and scope This dataset evaluates an LLM's reasoning ability in a legal context. Each question presents a realistic scenario involving competing legal principals, and asks the LLM to present a correct legal resolution with sufficient justification based on precedent. The dataset was created using slip opinions of the US Supreme Court from the 2024 term, taken from the Supreme Court website. Dataset Creation Method The benchmark was created using RELAI’s data… See the full description on the dataset page: https://huggingface.co/datasets/relai-ai/legal-scenarios-SCOTUS-2024-decisions.textquestion-answeringn<1K9 likes122 downloads1y agoHugging Face06lianghsun /tw-legal-benchmark-v2 Taiwan Legal Benchmark v2 A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese, built from 15 years (2012–2026) of national examinations published by the Ministry of Examination (考選部). Supersedes tw-legal-benchmark-v1 (209 questions) with 17,002 deduplicated questions across 15 legal domains. Overview Property Value Questions 17,002 (deduplicated) Years 2012–2026 Source papers 1,040 official exam papers Format… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2.tabularquestion-answering10K<n<100K3 likes121 downloads1mo agoHugging Face07Taylor658 /synthetic-legal ⚖️ Synthetic Legal (Query, Response) Dataset 📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers. ⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.texttext-generation100K<n<1M10 likes105 downloads13d agoHugging Face08mtntasci /turkish-legal-rag Turkish Legal RAG Corpus — Türk Hukuku için Açık RAG Datasetı Tek cümle: 25 önemli Türk kanununun (mevzuat.gov.tr kaynaklı, madde bazlı temiz chunk'lar) + 290 manuel doğrulanmış soru-cevap altın benchmark'ının olduğu açık kaynak Türkçe hukuk RAG datasetı. 🇹🇷 Türkçe Özet — Bu dataset, Türkçe hukuk uygulamaları için sıfırdan üretilmiş açık ve denetlenebilir bir RAG corpus'udur. mevzuat.gov.tr üzerinden alınan 25 ana kanunun madde madde temizlenmiş, chunk'lanmış sürümünü (6.350… See the full description on the dataset page: https://huggingface.co/datasets/mtntasci/turkish-legal-rag.tabulartext-retrieval1K<n<10K2 likes78 downloads4mo agoHugging Face09ciol-research /multilevel-legal-reasoning Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai. 🧭 Purpose and Scope The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.tabulartext-generationn<1K7 likes65 downloads1y agoHugging Face10Advocate512 /legalbench Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/Advocate512/legalbench.tabulartext-classification10K<n<100K0 likes63 downloads5mo agoHugging Face11Jel1f1sh /tw-legal-benchmark-v2 Taiwan Legal Benchmark v2 A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese, built from 15 years (2012–2026) of national examinations published by the Ministry of Examination (考選部). Supersedes tw-legal-benchmark-v1 (209 questions) with 17,002 deduplicated questions across 15 legal domains. Overview Property Value Questions 17,002 (deduplicated) Years 2012–2026 Source papers 1,040 official exam papers Format… See the full description on the dataset page: https://huggingface.co/datasets/Jel1f1sh/tw-legal-benchmark-v2.tabularquestion-answering10K<n<100K0 likes60 downloads29d agoHugging Face12Arailym-tleubayeva /legalup-laws Kazakhstan Legal Acts Dataset (LegalUp) Dataset Summary The LegalUp dataset contains structured metadata for legislative documents of the Republic of Kazakhstan. The current release includes 392,084 legislative document records extracted from a PostgreSQL database. The dataset is designed for: Legal Retrieval-Augmented Generation (Legal RAG) Information Retrieval Legal Search Question Answering Semantic Search Legal NLP Benchmark Construction Academic Research… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/legalup-laws.textquestion-answering100K<n<1M0 likes49 downloads15d agoHugging Face13somosnlp /justicio_evaluacion_ideonidad_preguntas_legalesEste dataset nos permitirá evaluar la idoneidad de las preguntas generadas para su uso dentro de la plataforma Justicio, un archivero que permite consultar desde una interfaz chat las distintas legislaciones, tanto a nivel nacional derivadas del Boletín Oficial del Estado, así como de las derivadas de las distintas Comunidades Autónomas. Internamente, Justicio utiliza un esquema de tipo RAG (Retrieval-Augmented Generation) en el que se localizan aquellos fragmentos almacenados más similares a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/justicio_evaluacion_ideonidad_preguntas_legales.textquestion-answeringn<1K1 likes45 downloads3y agoHugging Face14SaiCharanChetpelly /mmlu-legal-dataset-mcqtextquestion-answering1K<n<10K1 likes45 downloads2y agoHugging Face15adalat-ai /Indian-Legal-Retrieval-Generationgated Indian-Legal-Retrieval-Generation An expert-verified evaluation set for retrieval-augmented question answering over Indian court / legal documents. This is the small benchmark used in CourtNav. Paper: CourtNav: Voice-Guided, Anchor-Accurate Navigation of Long Legal Documents in Courtrooms — Sai Khadloya, Kush Juvekar, Arghya Bhattacharya, Utkarsh Saxena. Status: work in progress — contents and structure may still evolve. Overview 21 lawyer-verified question/answer… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/Indian-Legal-Retrieval-Generation.documentquestion-answeringn<1K0 likes33 downloads5mo agoHugging Face16roslein /Legal_advice_czech Legal Advice Dataset Dataset Description This dataset contains scraped legal questions (but also a simply informational content without any question present) and answers from Bezplatná Právní Poradna. The data consists of legal inquiries submitted by users and expert responses provided on the website. It is structured for ease of use in natural language processing (NLP) tasks related to legal text classification, question-answering models, and text summarization. However… See the full description on the dataset page: https://huggingface.co/datasets/roslein/Legal_advice_czech.textquestion-answering10K<n<100K0 likes29 downloads2y agoHugging Face17l0rdkr0n0s /albanian_legal_questions_answerstextquestion-answeringn<1K0 likes28 downloads1y agoHugging Face18gurumurthy3 /Legal-FAQ Legal FAQ Dataset Dataset Description The Legal FAQ Dataset is a collection of frequently asked legal questions along with their respective answers. This dataset is particularly useful for building question-answering systems, legal chatbots, or other applications in the domain of law and justice. Dataset Details Dataset ID: gurumurthy3/Legal-FAQ Total Entries: 2091 Columns: Question: A legal-related question. Answer: The corresponding legal response.… See the full description on the dataset page: https://huggingface.co/datasets/gurumurthy3/Legal-FAQ.textquestion-answering1K<n<10K0 likes25 downloads2y agoHugging Face19DanishMahdi /Sindhi_Legal_1 Dataset Description The Sindhi Legal Dataset is a parallel text dataset designed for legal NLP tasks in low-resource Sindhi language settings. It contains Pakistani legal QnA in Sindhi (Arabic script). The dataset is intended for applications such as legal text generation, machine translation, retrieval-augmented generation (RAG), and legal question answering systems. Each record consists of an input legal text segment and its corresponding output legal text segment.… See the full description on the dataset page: https://huggingface.co/datasets/DanishMahdi/Sindhi_Legal_1.textquestion-answering1K<n<10K1 likes19 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.