CoolFace
Datasetpublic

oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02

PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2 This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments. It is not the canonical text corpus. Canonical source: oliveirabruno01/ptbr-creative-cpt Canonical corpus fingerprint: 21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840 Identity Model/tokenizer: Qwen/Qwen3.5-0.8B-Base Context length: 2048 Data-prep version: v0.2 Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes68downloads
Dataset Card

PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2

This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments.

It is not the canonical text corpus.

Canonical source:

oliveirabruno01/ptbr-creative-cpt

Canonical corpus fingerprint:

21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840

Identity

  • —Model/tokenizer: Qwen/Qwen3.5-0.8B-Base
  • —Context length: 2048
  • —Data-prep version: v0.2
  • —Primary split policy: work-disjoint, source × genre stratified by actual tokenizer token mass
  • —Primary eval target: 8% of usable tokens
  • —Actual eval token fraction: 8.0086%
  • —Train units: 1,230
  • —Eval units: 122
  • —Train tokens: 20,428,414
  • —Eval tokens: 1,778,470
  • —Author-OOD diagnostic units: 18
  • —Decontamination exclusions: 0
  • —Era axis: disabled

QA policy

Human review is authoritative for exclusion decisions.

Jev/model audit scores were used only for diagnostics and review prioritization and never automatically exclude data.

Localized OCR artifacts in otherwise usable literary documents are retained.

Primary eval distribution

SplitSourceToken share
evalbackbone_books90.44%
evalFE-Unicamp5.00%
evalOJS2.91%
evalFCRB1.14%
evalOMP/Open Books0.51%
trainbackbone_books90.50%
trainFE-Unicamp4.10%
trainOJS3.03%
trainFCRB1.14%
trainOMP/Open Books1.07%
SplitGenreToken share
evalfiction70.70%
evaldrama10.01%
evalpoetry_cordel8.53%
evalmemoir_cronica7.21%
evalliterary_essay3.55%
trainfiction70.30%
traindrama10.16%
trainpoetry_cordel8.51%
trainmemoir_cronica7.37%
trainliterary_essay3.50%

The primary heldout set is intended for CPT transfer/loss evaluation. A smaller strict author-OOD subset is retained separately as a diagnostic.

Decontamination

After freezing the v0.2 heldout split, train → eval contamination checks found no removals under the configured exact normalized text, substantial paragraph-overlap, and MinHash near-duplicate criteria.

See:

  • —train_vs_eval_decontam_report.parquet
  • —split_manifest.parquet
  • —train_manifest.parquet
  • —eval_manifest.parquet

Token representation

token_store/unit_tokens.uint32.bin

contains concatenated token IDs for retained logical units.

token_store/unit_token_index_v02.parquet

maps each logical corpus unit to its token offset and length and to the corrected v0.2 split.

The token store was created once and reused when the v0.2 split was repaired; the corpus was not retokenized.

Packed data

packed/

contains context-2048 packed binary streams for:

  • —full train
  • —full primary eval
  • —source-family eval slices
  • —genre eval slices
  • —author-OOD diagnostic eval

See packed/pack_index.json for exact paths and block counts.

Era metadata

The historical/modern experimental axis is intentionally disabled in this artifact.

The canonical year field was not sufficiently reliable as a representation of the historical era of the underlying literary text, so no historical-vs-modern claim should be derived from it.

Reproducibility

Start with:

DATA_PREP_HANDOFF_v02.json

It records the canonical identity, split policy, tokenizer identity, data-prep identity, token-store paths, packing configuration, and next-phase handoff.

SHA256SUMS.json contains SHA-256 digests and byte sizes for all published artifact files.

Rights / provenance

This repository stores derived tokenized artifacts and metadata from the canonical corpus.

For source-level provenance, rights evidence, licenses, and text metadata, consult the canonical repository:

oliveirabruno01/ptbr-creative-cpt