nexoneAB/swedish-legal-decisions-raw-v1
Swedish Court Decisions — Svenska Domstolsavgöranden 55,096 court decisions spanning 45 years of Swedish case law, purpose-built for LLM training. The most comprehensive open dataset of Swedish appellate court decisions available for AI development. Sourced directly from the official Swedish Courts case law database via their public REST API and preprocessed into three ready-to-use training configurations. Why This Dataset Scale and depth: 55,096 decisions… See the full description on the dataset page: https://huggingface.co/datasets/nexoneAB/swedish-legal-decisions-raw-v1.
Swedish Court Decisions — Svenska Domstolsavgöranden
55,096 court decisions spanning 45 years of Swedish case law, purpose-built for LLM training.
The most comprehensive open dataset of Swedish appellate court decisions available for AI development. Sourced directly from the official Swedish Courts case law database via their public REST API and preprocessed into three ready-to-use training configurations.
Why This Dataset
- Scale and depth: 55,096 decisions covering 1981 to present — the full sweep of modern Swedish jurisprudence
- Precedent-aware: Every record is flagged with
ar_vagledande(is_precedent), enabling targeted training on landmark rulings that shape Swedish law - Tri-modal design: Three subsets serve distinct training objectives — raw pretraining, instruction tuning, and structured information extraction
- Production-ready: Cleaned full text, consistent schemas, train/validation/test splits in Parquet format
- Open provenance: CC0 source data from Domstolsverket (Swedish Courts Administration), published on dataportal.se
Courts Covered
Decisions from all major Swedish appellate courts:
Dataset Configurations
pretrain — Continued Pre-training
Full-text decisions for domain-adaptive pretraining. Ideal for adapting a general Swedish or multilingual LLM to the legal domain.
Tip: Filter on ar_vagledande == True to extract the subset of landmark, precedent-setting rulings — essential for training models that reason about legal authority.instruct — Instruction Tuning / SFT
Structured prompt–response pairs for supervised fine-tuning. Five legally meaningful task types, written in Swedish, cover the core competencies of a legal assistant.
Task types (`uppgift`):
This configuration is ready to drop into any SFT pipeline (e.g. TRL, LLaMA-Factory, Axolotl) with minimal preprocessing.
structured — Legal Information Extraction
Full text paired with JSON-structured section annotations, suitable for fine-tuning models on document parsing, section classification, and legal entity extraction.
Sections in `sektioner_json`:
Quick Start
from datasets import load_dataset
# Full-text pretraining corpus
ds = load_dataset("your-org/svenska-domstolsavgoranden", "pretrain")
# Precedent decisions only
precedents = ds["train"].filter(lambda x: x["ar_vagledande"])
# Instruction tuning
sft = load_dataset("your-org/svenska-domstolsavgoranden", "instruct")
# Structured extraction
structured = load_dataset("your-org/svenska-domstolsavgoranden", "structured")Data Source & Coverage
Ethics & Privacy
- All decisions are public records published by Domstolsverket under CC0
- Personal data in published decisions is partially anonymized by the courts prior to publication
- This dataset is intended for research and model training purposes
- Users are responsible for ensuring GDPR compliance in downstream applications
- Not intended for re-identification of individuals
Citation
@dataset{swedish_court_decisions,
title = {Swedish Court Decisions — Svenska Domstolsavgöranden},
source = {domstol.se (Domstolsverket)},
url = {https://rattspraxis.etjanst.domstol.se},
license = {CC0},
records = {55096},
years = {1981--2025},
}Legal Notice
This dataset contains court decisions published openly by Domstolsverket (Swedish Courts Administration) under CC0 license via their public REST API. Personal data has been anonymized by Domstolsverket prior to publication. Users are responsible for ensuring GDPR compliance in their own use cases. Not intended for re-identification of individuals.
