CoolFace
Datasetpublic

joecwales/whiteglove-legal-2026

WhiteGlove Legal Knowledge Corpus Law StackExchange + Project Gutenberg — 2026 Pipeline: WhiteGlove Spectral Curation | Domain: Legal | Sources: Law StackExchange · Project Gutenberg LCC-K Dataset Summary A clean, deduplicated legal knowledge corpus combining two complementary sources: Law StackExchange Q&A (statute citations in context) and Project Gutenberg public domain legal treatises (Blackstone, International Law, Constitutional history… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-legal-2026.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes14downloads
Dataset Card

WhiteGlove Legal Knowledge Corpus

Law StackExchange + Project Gutenberg — 2026

Pipeline: WhiteGlove Spectral Curation | Domain: Legal | Sources: Law StackExchange · Project Gutenberg LCC-K


Dataset Summary

A clean, deduplicated legal knowledge corpus combining two complementary sources: Law StackExchange Q&A (statute citations in context) and Project Gutenberg public domain legal treatises (Blackstone, International Law, Constitutional history, Nuremberg transcripts). Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication and semantic chunking.

MetricLaw StackExchangeGutenberg LawCombined
Raw chunks processed60,27043,628103,898
After SimHash-128 dedup21,7147,55329,267
Deduplication rate63.9%82.7%71.8%
Train records20,6207,17727,797
Validation records1,0943761,470
Avg quality score0.57540.5536~0.57
Avg words per record~490~490~490
Estimated tokens~9.9M~4.9M~14.8M
LicenseCC BY-SA 4.0Public DomainSee below

Sources

Law StackExchange (law.stackexchange.com_en_all_2026-02)

Real legal Q&A with statute citations in context. UCC, contract law, property, criminal procedure, jurisdiction, constitutional questions answered with references to actual code sections. The Q&A format produces naturally high-quality training signal — questions establish context, answers provide doctrine with citation.

Coverage includes:

  • —UCC Articles 1-9 (negotiable instruments, secured transactions, sales)
  • —Contract formation, breach, remedies, damages
  • —Constitutional law (First, Fourth, Fifth Amendment)
  • —Criminal procedure and evidence
  • —Property law, landlord/tenant, easements
  • —Employment law, discrimination, wrongful termination
  • —Immigration and citizenship
  • —International law and jurisdiction

License: CC BY-SA 4.0 — requires attribution and share-alike.


Project Gutenberg — LCC-K (Law) (gutenberg_en_lcc-k_2026-03)

Public domain legal treatises. High dedup rate (82.7%) reflects the internal repetition in multi-volume works — all 18 Nuremberg trial volumes, multiple editions of the same texts — the pipeline strips the noise and keeps the doctrine.

Notable works included:

  • —Blackstone's Commentaries on the Laws of England (Book I)
  • —Oppenheim's International Law: A Treatise (Vol. 1 & 2)
  • —The Constitution of the United States: Analysis and Interpretation
  • —History of the Origin, Formation, and Adoption of the Constitution (Vol. 1 & 2)
  • —Trial of the Major War Criminals — Nuremberg Military Tribunal (18 volumes)
  • —Magna Carta: A Commentary on the Great Charter of King John
  • —Our Legal Heritage: King AEthelbert — King George III
  • —The Rights of War and Peace (Grotius)
  • —The Visigothic Code (Forum Judicum)
  • —Medical Jurisprudence and Toxicology (Vol. 1-3)
  • —Tribal Custom in Anglo-Saxon Law
  • —Essays on the Constitution of the United States

License: Public domain (pre-1928 publication or US government work).


Pipeline Architecture

Law StackExchange .zim (176MB)          Gutenberg LCC-K .zim (235MB)
         │                                        │
         ▼ zim_extractor.py                       ▼ zim_extractor.py
  55,893 raw HTML shards               2,342 entries (554 extracted)
         │                                        │
         ▼ rechunk_domain.py                      ▼ rechunk_domain.py
  500-word semantic chunks             500-word semantic chunks
  60,270 shards                        43,628 shards
         │                                        │
         └──────────────┬─────────────────────────┘
                        ▼ export_training_set.py
          SimHash-64 deduplication (Hamming ratio < 0.12)
          Quality scoring (unique word ratio, punctuation density)
          Train/validation split (95/5)
                        │
                        ▼
         HF-compatible JSONL (train.jsonl, validation.jsonl)
         + manifest.json (full provenance per source)

Key properties:

  • —Two-source fusion — Q&A context from Law SE + foundational doctrine from Gutenberg
  • —Zero cloud dependency — entire pipeline runs air-gapped on consumer hardware
  • —SimHash-128 deduplication — catches near-duplicates across both sources
  • —Semantic chunking — 500-word windows preserve topical coherence
  • —Full provenance — every record traces back to source, source_tag, chunk index
  • —Source-filterable — source field identifies Law StackExchange vs Project Gutenberg Law

Record Schema

json
{
  "text": "...",
  "instruction": "Provide accurate information about: {title}",
  "output": "...",
  "id": "law_chunk_000042",
  "source": "Law StackExchange",
  "source_tag": "LawSE-2026-02",
  "domain": "legal",
  "license": "public-domain",
  "title": "Is verbal contract enforceable?",
  "path": "law.stackexchange.com/questions/...",
  "source_shard_id": "law_012345",
  "chunk_index": 0,
  "total_chunks": 2,
  "word_count": 487,
  "quality_score": 0.6102,
  "simhash": 9823741029384,
  "schema_version": "whiteglove-v1"
}

Filter by source:

python
# Law SE only
law_se = ds["train"].filter(lambda x: x["source"] == "Law StackExchange")

# Gutenberg treatises only
treatises = ds["train"].filter(lambda x: x["source"] == "Project Gutenberg Law")

# High quality only
high_q = ds["train"].filter(lambda x: x["quality_score"] > 0.65)

Usage

python
from datasets import load_dataset

ds = load_dataset("joecwales/whiteglove-legal-2026")
print(ds["train"][0]["text"])

# Filter by source
law_se = ds["train"].filter(lambda x: x["source"] == "Law StackExchange")
treatises = ds["train"].filter(lambda x: x["source"] == "Project Gutenberg Law")

Fine-tuning (SFT format)

python
# Compatible with TRL SFTTrainer, LLaMA Factory, Axolotl
ds = load_dataset("joecwales/whiteglove-legal-2026")
# instruction/output fields pre-formatted for supervised fine-tuning

Reproduce It

bash
# 1. Download ZIMs from kiwix.org
curl -L -o law_qa.zim "https://lbo.download.kiwix.org/zim/stack_exchange/law.stackexchange.com_en_all_2026-02.zim"
curl -L -o gutenberg_law.zim "https://lbo.download.kiwix.org/zim/gutenberg/gutenberg_en_lcc-k_2026-03.zim"

# 2. Extract
python3 zim_extractor.py --zim law_qa.zim --out shards/legal_qa --prefix law --source "Law StackExchange"
python3 zim_extractor.py --zim gutenberg_law.zim --out shards/gutenberg_law --prefix gut --source "Project Gutenberg Law"

# 3. Rechunk
python3 rechunk_domain.py --staging shards/legal_qa --out shards/legal_qa_chunked --prefix law_chunk --source "Law StackExchange" --domain legal --tag "LawSE-2026-02"
python3 rechunk_domain.py --staging shards/gutenberg_law --out shards/gutenberg_law_chunked --prefix gut_chunk --source "Project Gutenberg Law" --domain legal --tag "Gutenberg-LCC-K-2026-03"

# 4. Export (run separately, combine manually or merge JSONL files)
python3 export_training_set.py --shards shards/legal_qa_chunked --out exports/legal_qa --domain legal --tag "LawSE-2026-02" --source "Law StackExchange" --glob "law_chunk_*.json"
python3 export_training_set.py --shards shards/gutenberg_law_chunked --out exports/gutenberg_law --domain legal --tag "Gutenberg-LCC-K-2026-03" --source "Project Gutenberg Law" --glob "gut_chunk_*.json"

WhiteGlove Pipeline

This dataset was produced by the WhiteGlove Spectral Curation Pipeline — a self-contained, air-gapped dataset factory. Core properties:

  • —SimHash-128 deduplication (O(k) lookup, no vector database required)
  • —Semantic landmark indexing for Faith-Less retrieval (238ms over 33k shards, offline)
  • —Time-aware chunking for dynamic corpora (game code, versioned docs, event sequences)
  • —Sector-agnostic: swap the ZIM, same pipeline, new domain
Embed once. Query forever. No cloud. No fine-tuning required.

License

SourceLicenseRequirements
Law StackExchangeCC BY-SA 4.0Attribution + share-alike
Project Gutenberg LCC-KPublic DomainNone
Pipeline scriptsMITNone

Law StackExchange attribution: Content sourced from law.stackexchange.com, licensed under CC BY-SA 4.0. Original authors retain copyright. This dataset is a derivative work under the same license.


Citation

bibtex
@dataset{whiteglove_legal_2026,
  title     = {WhiteGlove Legal Knowledge Corpus — Law SE + Gutenberg 2026},
  author    = {Wales, Joe},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/joecwales/whiteglove-legal-2026},
  note      = {Produced by the WhiteGlove Spectral Curation Pipeline.
               Sources: Law StackExchange (CC BY-SA 4.0), Project Gutenberg LCC-K (Public Domain).}
}