interpretive-canons
interpretive-canons-eval-runs
Interpretive Canons — evaluation runs
Companion release to the paper Classifying Interpretive Canons at the Sentence
Level: A Benchmark from the German Federal Constitutional Court. This repository
holds the reproducibility artifacts behind the paper's results: the raw model
predictions for every reported cell, the LLM judge's recorded decisions for the
statutory-reference subtask, and the exact prompts that produced the runs.
It is the third of three companion repositories:… See the full description on the dataset page: https://huggingface.co/datasets/felix453/interpretive-canons-eval-runs.interpretive-canons-instances
Interpretive Canons — instance-level
Companion benchmark to the paper Classifying Interpretive Canons at the
Sentence Level: A Benchmark from the German Federal Constitutional Court. This
is the instance-level release: one record per (subtask, candidate) pair,
grouped into eight subtasks, each with decision-disjoint train/validation/test
splits (no decision appears in more than one split). The companion
decision-level release (one record per exhaustively annotated decision, for… See the full description on the dataset page: https://huggingface.co/datasets/felix453/interpretive-canons-instances.interpretive-canons-raw-export
Classifying Interpretive Canons — Raw Label Studio Export (pre-merge snapshot)
Raw, pre-merge Label Studio annotation export (project 46, snapshot 2026-07-25).
Two fields needed to reconstruct the konkretes_gesetz (statutory-reference)
subtask survive only in this pre-merge form: the per-reading
determinate-content judgment bestimmt_moeglicheDeutung, which the
postprocessing merge drops entirely, and the list of concrete provisions, which
the merge flattens into a single… See the full description on the dataset page: https://huggingface.co/datasets/felix453/interpretive-canons-raw-export.interpretive-canons-decisions
Interpretive Canons — decision-level
Companion benchmark to the paper Classifying Interpretive Canons at the
Sentence Level: A Benchmark from the German Federal Constitutional Court. This
is the decision-level release: one record per fully (exhaustively) annotated
decision, intended for end-to-end evaluation that runs the full pipeline over a
whole decision. Only the 15 exhaustively annotated decisions are included; the
selectively annotated decisions are not, because their… See the full description on the dataset page: https://huggingface.co/datasets/felix453/interpretive-canons-decisions.
