joecwales/whiteglove-legal-2026
WhiteGlove Legal Knowledge Corpus Law StackExchange + Project Gutenberg — 2026 Pipeline: WhiteGlove Spectral Curation | Domain: Legal | Sources: Law StackExchange · Project Gutenberg LCC-K Dataset Summary A clean, deduplicated legal knowledge corpus combining two complementary sources: Law StackExchange Q&A (statute citations in context) and Project Gutenberg public domain legal treatises (Blackstone, International Law, Constitutional history… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-legal-2026.
WhiteGlove Legal Knowledge Corpus
Law StackExchange + Project Gutenberg — 2026
Pipeline: WhiteGlove Spectral Curation | Domain: Legal | Sources: Law StackExchange · Project Gutenberg LCC-K
Dataset Summary
A clean, deduplicated legal knowledge corpus combining two complementary sources: Law StackExchange Q&A (statute citations in context) and Project Gutenberg public domain legal treatises (Blackstone, International Law, Constitutional history, Nuremberg transcripts). Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication and semantic chunking.
Sources
Law StackExchange (law.stackexchange.com_en_all_2026-02)
Real legal Q&A with statute citations in context. UCC, contract law, property, criminal procedure, jurisdiction, constitutional questions answered with references to actual code sections. The Q&A format produces naturally high-quality training signal — questions establish context, answers provide doctrine with citation.
Coverage includes:
- UCC Articles 1-9 (negotiable instruments, secured transactions, sales)
- Contract formation, breach, remedies, damages
- Constitutional law (First, Fourth, Fifth Amendment)
- Criminal procedure and evidence
- Property law, landlord/tenant, easements
- Employment law, discrimination, wrongful termination
- Immigration and citizenship
- International law and jurisdiction
License: CC BY-SA 4.0 — requires attribution and share-alike.
Project Gutenberg — LCC-K (Law) (gutenberg_en_lcc-k_2026-03)
Public domain legal treatises. High dedup rate (82.7%) reflects the internal repetition in multi-volume works — all 18 Nuremberg trial volumes, multiple editions of the same texts — the pipeline strips the noise and keeps the doctrine.
Notable works included:
- Blackstone's Commentaries on the Laws of England (Book I)
- Oppenheim's International Law: A Treatise (Vol. 1 & 2)
- The Constitution of the United States: Analysis and Interpretation
- History of the Origin, Formation, and Adoption of the Constitution (Vol. 1 & 2)
- Trial of the Major War Criminals — Nuremberg Military Tribunal (18 volumes)
- Magna Carta: A Commentary on the Great Charter of King John
- Our Legal Heritage: King AEthelbert — King George III
- The Rights of War and Peace (Grotius)
- The Visigothic Code (Forum Judicum)
- Medical Jurisprudence and Toxicology (Vol. 1-3)
- Tribal Custom in Anglo-Saxon Law
- Essays on the Constitution of the United States
License: Public domain (pre-1928 publication or US government work).
Pipeline Architecture
Law StackExchange .zim (176MB) Gutenberg LCC-K .zim (235MB)
│ │
▼ zim_extractor.py ▼ zim_extractor.py
55,893 raw HTML shards 2,342 entries (554 extracted)
│ │
▼ rechunk_domain.py ▼ rechunk_domain.py
500-word semantic chunks 500-word semantic chunks
60,270 shards 43,628 shards
│ │
└──────────────┬─────────────────────────┘
▼ export_training_set.py
SimHash-64 deduplication (Hamming ratio < 0.12)
Quality scoring (unique word ratio, punctuation density)
Train/validation split (95/5)
│
▼
HF-compatible JSONL (train.jsonl, validation.jsonl)
+ manifest.json (full provenance per source)Key properties:
- Two-source fusion — Q&A context from Law SE + foundational doctrine from Gutenberg
- Zero cloud dependency — entire pipeline runs air-gapped on consumer hardware
- SimHash-128 deduplication — catches near-duplicates across both sources
- Semantic chunking — 500-word windows preserve topical coherence
- Full provenance — every record traces back to source,
source_tag, chunk index - Source-filterable —
sourcefield identifies Law StackExchange vs Project Gutenberg Law
Record Schema
{
"text": "...",
"instruction": "Provide accurate information about: {title}",
"output": "...",
"id": "law_chunk_000042",
"source": "Law StackExchange",
"source_tag": "LawSE-2026-02",
"domain": "legal",
"license": "public-domain",
"title": "Is verbal contract enforceable?",
"path": "law.stackexchange.com/questions/...",
"source_shard_id": "law_012345",
"chunk_index": 0,
"total_chunks": 2,
"word_count": 487,
"quality_score": 0.6102,
"simhash": 9823741029384,
"schema_version": "whiteglove-v1"
}Filter by source:
# Law SE only
law_se = ds["train"].filter(lambda x: x["source"] == "Law StackExchange")
# Gutenberg treatises only
treatises = ds["train"].filter(lambda x: x["source"] == "Project Gutenberg Law")
# High quality only
high_q = ds["train"].filter(lambda x: x["quality_score"] > 0.65)Usage
from datasets import load_dataset
ds = load_dataset("joecwales/whiteglove-legal-2026")
print(ds["train"][0]["text"])
# Filter by source
law_se = ds["train"].filter(lambda x: x["source"] == "Law StackExchange")
treatises = ds["train"].filter(lambda x: x["source"] == "Project Gutenberg Law")Fine-tuning (SFT format)
# Compatible with TRL SFTTrainer, LLaMA Factory, Axolotl
ds = load_dataset("joecwales/whiteglove-legal-2026")
# instruction/output fields pre-formatted for supervised fine-tuningReproduce It
# 1. Download ZIMs from kiwix.org
curl -L -o law_qa.zim "https://lbo.download.kiwix.org/zim/stack_exchange/law.stackexchange.com_en_all_2026-02.zim"
curl -L -o gutenberg_law.zim "https://lbo.download.kiwix.org/zim/gutenberg/gutenberg_en_lcc-k_2026-03.zim"
# 2. Extract
python3 zim_extractor.py --zim law_qa.zim --out shards/legal_qa --prefix law --source "Law StackExchange"
python3 zim_extractor.py --zim gutenberg_law.zim --out shards/gutenberg_law --prefix gut --source "Project Gutenberg Law"
# 3. Rechunk
python3 rechunk_domain.py --staging shards/legal_qa --out shards/legal_qa_chunked --prefix law_chunk --source "Law StackExchange" --domain legal --tag "LawSE-2026-02"
python3 rechunk_domain.py --staging shards/gutenberg_law --out shards/gutenberg_law_chunked --prefix gut_chunk --source "Project Gutenberg Law" --domain legal --tag "Gutenberg-LCC-K-2026-03"
# 4. Export (run separately, combine manually or merge JSONL files)
python3 export_training_set.py --shards shards/legal_qa_chunked --out exports/legal_qa --domain legal --tag "LawSE-2026-02" --source "Law StackExchange" --glob "law_chunk_*.json"
python3 export_training_set.py --shards shards/gutenberg_law_chunked --out exports/gutenberg_law --domain legal --tag "Gutenberg-LCC-K-2026-03" --source "Project Gutenberg Law" --glob "gut_chunk_*.json"WhiteGlove Pipeline
This dataset was produced by the WhiteGlove Spectral Curation Pipeline — a self-contained, air-gapped dataset factory. Core properties:
- SimHash-128 deduplication (O(k) lookup, no vector database required)
- Semantic landmark indexing for Faith-Less retrieval (238ms over 33k shards, offline)
- Time-aware chunking for dynamic corpora (game code, versioned docs, event sequences)
- Sector-agnostic: swap the ZIM, same pipeline, new domain
Embed once. Query forever. No cloud. No fine-tuning required.
License
Law StackExchange attribution: Content sourced from law.stackexchange.com, licensed under CC BY-SA 4.0. Original authors retain copyright. This dataset is a derivative work under the same license.
Citation
@dataset{whiteglove_legal_2026,
title = {WhiteGlove Legal Knowledge Corpus — Law SE + Gutenberg 2026},
author = {Wales, Joe},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/joecwales/whiteglove-legal-2026},
note = {Produced by the WhiteGlove Spectral Curation Pipeline.
Sources: Law StackExchange (CC BY-SA 4.0), Project Gutenberg LCC-K (Public Domain).}
}