datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alloy-sovereign-eval-runs
Alloy Sovereign Eval Runs · the honest first measured run
Append-only measured eval runs produced by routing SZL's
K-Verify Benchmark v1
through the live Alloy governed-inference stack on SZL's own sovereign
metal (provider: sovereign, zero cloud, zero spend). Each row is one
inference: its verdict, latency, NVML-measured energy, and a
signed receipt id that is re-checkable against the live Alloy receipt chain.
Built and maintained by SZL Holdings. Apache-2.0.… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/alloy-sovereign-eval-runs.interpretive-canons-eval-runs
Interpretive Canons — evaluation runs
Companion release to the paper Classifying Interpretive Canons at the Sentence
Level: A Benchmark from the German Federal Constitutional Court. This repository
holds the reproducibility artifacts behind the paper's results: the raw model
predictions for every reported cell, the LLM judge's recorded decisions for the
statutory-reference subtask, and the exact prompts that produced the runs.
It is the third of three companion repositories:… See the full description on the dataset page: https://huggingface.co/datasets/felix453/interpretive-canons-eval-runs.eval-arena-runs
