CoolFace
Datasetpublic

tardellirs/brazilian-federal-law-corpus

Brazilian Federal Law Corpus (PT-BR) 41,526 passages of Brazilian federal legislation (laws and decrees — articles, titles, ementas) scraped from planalto.gov.br. Official Brazilian legal texts are public domain (Lei 9.610/1998, art. 8). A clean corpus for Portuguese legal IR / RAG / language modeling, with no LLM-generated content — pure law text — and decontaminated against MTEB(por). Companion to the synthetic QA pairs in tardellirs/brazilian-legal-tax-qa-synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-federal-law-corpus.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
4likes89downloads
Dataset Card

Brazilian Federal Law Corpus (PT-BR)

41,526 passages of Brazilian federal legislation (laws and decrees — articles, titles, ementas) scraped from planalto.gov.br. Official Brazilian legal texts are public domain (Lei 9.610/1998, art. 8). A clean corpus for Portuguese legal IR / RAG / language modeling, with no LLM-generated content — pure law text — and decontaminated against MTEB(por).

Companion to the synthetic QA pairs in `tardellirs/brazilian-legal-tax-qa-synthetic`.

Decontamination

Every passage sharing >20% of its 8-gram shingles (boilerplate auto-removed) with any MTEB(por) test corpus (JurisTCU, BRTaxQAR, FaQuAD, MedPT, Quati) was removed — 14,936 of 56,462 (26.5%) dropped → 41,526 kept with 0% exact and <20% content overlap, so it is safe to train on and evaluate on those benchmarks.

Fields

fielddescription
passagethe Brazilian federal-law passage
leilaw/decree identifier
anoyear
ementaofficial summary
is_fiscalwhether the law is tax/fiscal

Usage

python
from datasets import load_dataset
ds = load_dataset("tardellirs/brazilian-federal-law-corpus", split="train")

License & provenance

CC-BY-4.0 (the source legal texts are public domain; scraped from planalto.gov.br). Benchmark used for decontamination: MTEB(por) · leaderboard.