scotus
Datasets
All datasets matching “scotus”scotus-vocal-data
Mimir: Supreme Court voice and vote data
Mimir studies whether measurements of justices' speech during oral argument help predict
their votes. Start with the evaluated model and its model card:
Mimir vote predictor:
the selected model, portable inference, final evaluation, coverage and limitations.
Complete supporting model data:
fitting data, separate final inputs and targets, candidate models, predictions, acoustic
measurements, new word and diarization outputs, source… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/scotus-vocal-data.scotus-opinionsscotus_opinions
SCOTUS Opinions
Collection of embeddings of SCOTUS opinions from argument years 1991 to 2024 (as of Jan 2025). Source material is available from the official SCOTUS web site.
Data is generated using text-embedding-ada-002, with a chunk size of 2560 tokens and 256-token overlap.
Usage
To load the dataset into langchain VectorStore:
vectorstore = SKLearnVectorStore(
embedding=OpenAIEmbeddings(openai_api_key=api_key),
persist_path=$PATH… See the full description on the dataset page: https://huggingface.co/datasets/chowalex/scotus_opinions.SCOTUS-Dec
Polish SCOTUS-Dec
Polish SCOTUS-Dec is a long-document legal text classification dataset
derived from the publicly available
SCOTUS dataset.
The dataset contains Polish translations of U.S. Supreme Court opinions.
The task is a binary document classification problem in which the goal
is to predict the ideological direction associated with the court decision:
liberal or conservative.
SCOTUS-Dec is part of the LongContext benchmark introduced with
Polish ModernBERT.… See the full description on the dataset page: https://huggingface.co/datasets/mmichall/SCOTUS-Dec.oral-arguments-scotus
[!NOTE]
Dataset origin: https://www.ortolang.fr/market/corpora/oral-arguments-scotus
[!CAUTION]
Vous devez vous rendre sur le site d'Ortholang et vous connecter afin de télécharger l'intégralité des données (seul un sous-échantillon est disponible dans ce répertoire).
Description
Ce corpus comprend 112 sessions d’oral arguments – une session correspondant à une affaire et de façon générale à deux oral arguments et parfois un rebuttal argument – s’étalant du 9 octobre 2018… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/oral-arguments-scotus.SCOTUS-Dom
Polish SCOTUS-Dom
Polish SCOTUS-Dom is a long-document legal text classification dataset
derived from the publicly available
SCOTUS dataset.
The dataset contains Polish translations of U.S. Supreme Court opinions.
The task is an 11-class document classification problem in which the goal
is to predict the legal issue area of a court case.
SCOTUS-Dom is part of the LongContext benchmark introduced with
Polish ModernBERT.
Dataset statistics
Split
Examples… See the full description on the dataset page: https://huggingface.co/datasets/mmichall/SCOTUS-Dom.
