JuDDGES/pl-court-raw
Dataset Card for JuDDGES/pl-court-raw Dataset Summary The dataset consists of Polish Court judgments available at https://orzeczenia.ms.gov.pl/, containing full content of the judgments along with metadata sourced from official API and extracted from the judgment contents. This dataset contains raw data. For instruction dataset see JuDDGES/pl-court-instruct. For graph dataset see JuDDGES/pl-court-graph. Supported Tasks and Leaderboards The dataset… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/pl-court-raw.
Dataset Card for JuDDGES/pl-court-raw
Table of Contents
- Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
- Statistics
Dataset Description
- Homepage: TBA
- Repository: https://github.com/pwr-ai/JuDDGES
- Paper: TBA
- Point of Contact: lukasz.augustyniak@pwr.edu.pl; jakub.binkowski@pwr.edu.pl; albert.sawczyn@pwr.edu.pl
Dataset Summary
The dataset consists of Polish Court judgments available at https://orzeczenia.ms.gov.pl/, containing full content of the judgments along with metadata sourced from official API and extracted from the judgment contents. This dataset contains raw data. For instruction dataset see `JuDDGES/pl-court-instruct`. For graph dataset see `JuDDGES/pl-court-graph`.
Supported Tasks and Leaderboards
The dataset can be used for various tasks. However, it contains raw data acquired from official API, and we rather recommend using instruction dataset `JuDDGES/pl-court-instruct` for straightforward usage.
Languages
pl-PL Polish
Dataset Structure
Data Fields
Data Splits
This dataset is not split into subsets. The dataset has only train split.
Dataset Creation
For details on the dataset creation, see the paper [TBA]() and the code repository here.
Curation Rationale
Created to enable cross-jurisdictional legal analytics.
Source Data
Initial Data Collection and Normalization
- Download judgments metadata.
- Download judgments text (XML content of judgments).
- Download additional details available for each judgment.
- Map id of courts and departments to court name.
- Extract raw text from XML content and details of judgments not available through API.
- For further processing prepare local dataset dump in parquet file, version it with dvc and push to remote storage.
Who are the source language producers?
Produced by human legal professionals (judges, court clerks). Demographics was not analysed. Sourced from public court databases.
Annotations
Annotation process
No annotation was performed by us. All features were provided via API (anonymization and publication of the data were performed by court employees).
Who are the annotators?
As above.
Personal and Sensitive Information
Pseudoanonymized to comply with GDPR (art. 4 sec. 5 GDPR).
Considerations for Using the Data
Social Impact of Dataset
[More Information Needed]
Discussion of Biases
[More Information Needed]
Other Known Limitations
[More Information Needed]
Additional Information
Dataset Curators
[More Information Needed]
Licensing Information
We license the actual packaging of these data under Attribution 4.0 International (CC BY 4.0) https://creativecommons.org/licenses/by/4.0/
Citation Information
TBA
Statistics
Dataset size
Dataset size: 437_450 examples
