datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
interpretive-canons-eval-runs
Interpretive Canons — evaluation runs
Companion release to the paper Classifying Interpretive Canons at the Sentence
Level: A Benchmark from the German Federal Constitutional Court. This repository
holds the reproducibility artifacts behind the paper's results: the raw model
predictions for every reported cell, the LLM judge's recorded decisions for the
statutory-reference subtask, and the exact prompts that produced the runs.
It is the third of three companion repositories:… See the full description on the dataset page: https://huggingface.co/datasets/felix453/interpretive-canons-eval-runs.eval-arena-runs
