datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
german-courts
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rusheeliyer/german-courts.EU_Court_Human_Rights_Decisions
European Court of Human Rights Decisions Dataset
This dataset contains 9,820 decisions from the European Court of Human Rights (ECHR) scraped from HUDOC, the official database of ECHR case law.
Data Usage
This dataset is valuable for:
Building legal vector databases for RAG (Retrieval Augmented Generation)
Training Large Language Models focused on human rights law
Creating synthetic legal datasets
Legal text analysis and research
NLP tasks in international human rights… See the full description on the dataset page: https://huggingface.co/datasets/roslein/EU_Court_Human_Rights_Decisions.court_opinions_filtered_under_25kcourt_opinions_filtered_full_sizeindian-court-judgements-and-its-summariesCZE_constitutional_court_decisions
Czech Constitutional Court Decisions Dataset
This dataset contains decisions from the Constitutional Court of the Czech Republic scraped from NALUS, the official database of Constitutional Court decisions.
Data Usage
This dataset can be utilized for:
Training language models on legal texts
Creating synthetic legal datasets
Building vector databases for Retrieval Augmented Generation (RAG)
Legal text analysis and research
NLP tasks focused on Czech legal domain… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_constitutional_court_decisions.EU_Court_Human_Rights_Decisions
European Court of Human Rights Decisions Dataset
This dataset contains 9,820 decisions from the European Court of Human Rights (ECHR) scraped from HUDOC, the official database of ECHR case law.
Data Usage
This dataset is valuable for:
Building legal vector databases for RAG (Retrieval Augmented Generation)
Training Large Language Models focused on human rights law
Creating synthetic legal datasets
Legal text analysis and research
NLP tasks in international human… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/EU_Court_Human_Rights_Decisions.lamus-roberts-court-legal-arguments
LAMUS: Roberts Court Legal Arguments (2005-2025)
The Current Supreme Court Era - Chief Justice John Roberts
📋 Dataset Description
This dataset contains 362,891 sentences from U.S. Supreme Court opinions during the Roberts Court era (2005-2025), automatically labeled with legal argument categories. This represents the current Supreme Court under Chief Justice John G. Roberts Jr.
Why Roberts Court?
The Roberts Court is particularly significant for… See the full description on the dataset page: https://huggingface.co/datasets/LavanyaPobbathi/lamus-roberts-court-legal-arguments.CZE_Supreme_Court_Decision
Czech Supreme Court Decisions Dataset
This dataset contains decisions from the Supreme Court of the Czech Republic scraped from their official collection database.
Data Usage
This dataset is ideal for:
Building legal vector databases for RAG (Retrieval Augmented Generation)
Training language models on Czech civil and criminal law
Creating synthetic legal datasets
Legal text analysis and research
NLP tasks focused on Czech judicial domain
Legal Status
The… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_Supreme_Court_Decision.us-public-tennis-courts
US Public Tennis Courts — 4,298 Locations Across 17 Metros (2026)
Geocoded public tennis court locations for 17 major US metros (Austin, Dallas-Fort Worth, Denver, Fort Lauderdale, Houston, Los Angeles, Miami, Orange County, Phoenix, Portland, San Diego, San Francisco, San Jose, Seattle, Tampa, Washington DC, West Palm Beach): 4,298 locations with latitude/longitude, court counts (11,527 individual courts where counted), lighting, practice walls, and surface where known. Derived… See the full description on the dataset page: https://huggingface.co/datasets/rrhagentbiz/us-public-tennis-courts.indian_supreme_court_judgements_en_ta
Indian Supreme Court Judgements Dataset (Sentence-Level, Translated to Tamil)
Overview
This dataset contains Indian Supreme Court judgements that have been split into sentences and translated into Tamil. The original judgements were sourced from the Indian Kanoon website. The dataset is useful for legal text processing, multilingual NLP tasks, and cross-lingual legal studies.
Data Processing Pipeline
Sentence Splitting:
Used pySBD (Python Sentence Boundary… See the full description on the dataset page: https://huggingface.co/datasets/Narenameme/indian_supreme_court_judgements_en_ta.CZE_supreme_administrative_court_decisions
Czech Supreme Administrative Court Decisions Dataset
This dataset contains decisions from the Supreme Administrative Court of the Czech Republic scraped from their official search interface.
Data Usage
This dataset can be utilized for:
Training language models on administrative law texts
Creating synthetic legal datasets
Building vector databases for Retrieval Augmented Generation (RAG)
Administrative law text analysis and research
NLP tasks focused on Czech… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_supreme_administrative_court_decisions.mteb_brazilian_court_decisionsbrazilian_court_civil_decisionsgreek-supreme-court-decisionslegal-court-bundle-exhibit-index-pagination-coherence-risk-v0.1What this dataset does
You receive
bundle index
exhibit list
pagination plan
actual contents
missing flags
duplicate flags
You decide
coherent
or
incoherent
Daily use
bundle QC before filing
missing exhibit detection
pagination mismatch detection
court_data_RAG_unsupbot-courteous-interactionsThis dataset contains a collection of variations in polite and courteous responses generated by a conversational AI. This dataset is designed to enhance natural language understanding and generation models, focusing on responses that convey gratitude, appreciation, and helpfulness. Each entry in the dataset pairs a user input with multiple variations of a polite, appreciative response, aiming to enrich conversational models with diverse ways of expressing politeness and support.
analysed_court_rulingsjanaab_supreme-court-speech_test_embeddingslegal-court-fees-coherence-distortion-v0.1What this dataset is
You receive
fee amount
waiver rule
claimant means profile
drop off and default pattern
admin friction
reform signals
You decide
Do fees distort access to adjudication
Answer
coherent
or
incoherent
Why this matters
Fee distortion predicts
pro se collapse
default growth
forced settlement
trust erosion
court-booking-dataindian-court-judgements-and-its-summaries
