tardellirs/brazilian-federal-law-corpus
Brazilian Federal Law Corpus (PT-BR) 41,526 passages of Brazilian federal legislation (laws and decrees — articles, titles, ementas) scraped from planalto.gov.br. Official Brazilian legal texts are public domain (Lei 9.610/1998, art. 8). A clean corpus for Portuguese legal IR / RAG / language modeling, with no LLM-generated content — pure law text — and decontaminated against MTEB(por). Companion to the synthetic QA pairs in tardellirs/brazilian-legal-tax-qa-synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-federal-law-corpus.
Brazilian Federal Law Corpus (PT-BR)
41,526 passages of Brazilian federal legislation (laws and decrees — articles, titles, ementas) scraped from planalto.gov.br. Official Brazilian legal texts are public domain (Lei 9.610/1998, art. 8). A clean corpus for Portuguese legal IR / RAG / language modeling, with no LLM-generated content — pure law text — and decontaminated against MTEB(por).
Companion to the synthetic QA pairs in `tardellirs/brazilian-legal-tax-qa-synthetic`.
Decontamination
Every passage sharing >20% of its 8-gram shingles (boilerplate auto-removed) with any MTEB(por) test corpus (JurisTCU, BRTaxQAR, FaQuAD, MedPT, Quati) was removed — 14,936 of 56,462 (26.5%) dropped → 41,526 kept with 0% exact and <20% content overlap, so it is safe to train on and evaluate on those benchmarks.
Fields
Usage
from datasets import load_dataset
ds = load_dataset("tardellirs/brazilian-federal-law-corpus", split="train")License & provenance
CC-BY-4.0 (the source legal texts are public domain; scraped from planalto.gov.br). Benchmark used for decontamination: MTEB(por) · leaderboard.
