SNTIQ-Team/lexgraph
Lexgraph dataset The data plane of Lexgraph — German legislation modelled as Laws as Git: a temporal, multi-authority event log with HEAD, commits, open/closed branches and evidence-bound merges (Bund / Bayern / EU; Länder records only after verification at the originating Landtag). Built 2026-07-19. Git is a navigation metaphor, not a substitute for legal status. Every row's official source and status controls whether it is current law, a pending branch or a documented… See the full description on the dataset page: https://huggingface.co/datasets/SNTIQ-Team/lexgraph.
Lexgraph dataset
The data plane of [Lexgraph](https://github.com/SNTIQ-Team/lexgraph) — German legislation modelled as Laws as Git: a temporal, multi-authority event log with HEAD, commits, open/closed branches and evidence-bound merges (Bund / Bayern / EU; Länder records only after verification at the originating Landtag). Built 2026-07-19.
Git is a navigation metaphor, not a substitute for legal status. Every row's official source and status controls whether it is current law, a pending branch or a documented implementation/applicability link.
Lexgraph is deliberately two-layered. The German migration, asylum, social-law and related practice corpus has current full text and change history where the official/retrieval sources support it. federal_catalog.jsonl provides every act listed in the official GII table of contents as discovery metadata only; eu_index.jsonl provides EU breadth: all in-force directives (including implementing/delegated directives) and basic regulations exposed by CELLAR, as metadata only. It is not presented as deep coverage of every EU instrument.
The independent federal archive is split by claim strength. An official_federal_state_observations.jsonl row says only what complete GII state Lexgraph retrieved on observed_at; that date is not silently treated as commencement. official_federal_state_transitions.jsonl contains Lexgraph's own old/new diff between two such states and likewise leaves effective_at empty. A row enters official_transition_reviews.jsonl only after the complete state pair is matched to the final, integrity-checked BGBl amending command and the exact DIP commencement clause. Full portable states, including every captured norm, are in official_federal_state_objects.jsonl and are identified by the SHA-256 of canonical uncompressed JSON.
retrospective_legal_intervals.jsonl adds a true bitemporal projection: legal validity (effective_from/effective_to) is independent from Lexgraph knowledge time (knowledge_from/knowledge_to). The 2023+ final-BGBl inventory is exported separately as retrospective_amendment_events.jsonl. An event with a publication/effective date is not thereby a reconstructed historical consolidated text; unresolved sub-article dates and missing state pairs remain explicit in retrospective_gaps.jsonl. The complete manifest and a portable indexed SQLite representation are included as artifacts.
verified_reconstructions.json is a separate, reviewed claim class. Its derived_verified bodies were computed by reversing cardinality-checked final BGBl commands from a complete official GII anchor and then replaying the commands forward to reproduce that anchor byte-for-byte. They are complete Lexgraph reconstructions. The artifact deliberately keeps source_exact: false because NeuRIS/GII did not supply those bytes as a historical snapshot. The referenced deterministic-gzip objects are retained under federal_states/objects/ and remain separate from the official GII manifest.
Buzer is not an input to these four files and no private Buzer snapshot or synopsis is shipped. It may be used outside the export as a non-authoritative human QA/deep-link cross-check; all published state bytes, diffs and legal-date evidence above are reproduced independently from official sources.
decisions.jsonl combines reviewed cases with a forward-cumulative, corpus-filtered import from the seven official federal Rechtsprechung-im-Internet feeds. Automated norm links are citations, not claims that the court interpreted every cited provision. The table is not a catalogue of all German case law.
neuris_archive.jsonl is the exact logical NeuRIS changelog ledger. It keeps metadata-only, tombstone and failed-capture observations so the archive does not hide gaps. Only rows with capture_status: captured may reference bytes; each referenced official ZIP/XML/HTML artifact is re-hashed and exported under neuris_objects/. Temporary downloads, failed partials and unreferenced local cache files are excluded. ELI point-in-time and manifestation components are source identifiers, not silently asserted legal-effective dates.
For Bavaria, official version rows are amendment metadata. Word-level history is exported only when archived official pages yield an unambiguous state transition, plus complete forward diffs between daily snapshots. Missing history is kept missing instead of reconstructed. In bayern_word_diffs.jsonl, date is the promulgation/event date while effective_date is the date on which the consolidated wording changes; data consumers doing time travel must prefer effective_date.
The public API can serialize the current complete act or one provision as Markdown. It can also check out an exact archived GII observation when the requested date exists in the state store. Historical output remains explicitly labelled partial whenever the available source only supplies amendment excerpts rather than a lossless consolidated snapshot; the dataset does not imply exact historical text where the source archive cannot prove it.
watched_procedures.json preserves the current DIP/EUR-Lex observations, official document chronology and embedded status-change history. Its analysis object keeps verified facts, deterministic inferences and a qualitative forecast separate; likelihood is not presented as a statistically calibrated probability. Terminal procedures remain archived but leave the active polling set. amendment_fates.json separates reviewed roles in a parliamentary document chain from the mechanical checks performed against the current consolidated corpus.
DIP-derived rows use the attribution Deutscher Bundestag/Bundesrat – DIP. Lexgraph extraction, ranking, annotations and forecasts are transformations, not source statements. DIP makes its source data available free of charge at dip.bundestag.de.
Every source and file-level rights regime is documented in the repository's docs/SOURCES.md and docs/RIGHTS.md; reproduce the current outputs with refresh.sh and tools/export_hf.py. Statutory texts are official works under § 5 UrhG. The dataset is intentionally marked license: other: SNTIQ's licence covers its original annotations and software, not third-party or official source material. See RIGHTS.md before reuse.
Built by [SNTIQ](https://sntiq.com/).
