oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02
PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2 This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments. It is not the canonical text corpus. Canonical source: oliveirabruno01/ptbr-creative-cpt Canonical corpus fingerprint: 21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840 Identity Model/tokenizer: Qwen/Qwen3.5-0.8B-Base Context length: 2048 Data-prep version: v0.2 Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.
PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2
This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments.
It is not the canonical text corpus.
Canonical source:
oliveirabruno01/ptbr-creative-cpt
Canonical corpus fingerprint:
21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840
Identity
- Model/tokenizer:
Qwen/Qwen3.5-0.8B-Base - Context length:
2048 - Data-prep version:
v0.2 - Primary split policy: work-disjoint, source × genre stratified by actual tokenizer token mass
- Primary eval target: 8% of usable tokens
- Actual eval token fraction:
8.0086% - Train units:
1,230 - Eval units:
122 - Train tokens:
20,428,414 - Eval tokens:
1,778,470 - Author-OOD diagnostic units:
18 - Decontamination exclusions:
0 - Era axis: disabled
QA policy
Human review is authoritative for exclusion decisions.
Jev/model audit scores were used only for diagnostics and review prioritization and never automatically exclude data.
Localized OCR artifacts in otherwise usable literary documents are retained.
Primary eval distribution
The primary heldout set is intended for CPT transfer/loss evaluation. A smaller strict author-OOD subset is retained separately as a diagnostic.
Decontamination
After freezing the v0.2 heldout split, train → eval contamination checks found no removals under the configured exact normalized text, substantial paragraph-overlap, and MinHash near-duplicate criteria.
See:
train_vs_eval_decontam_report.parquetsplit_manifest.parquettrain_manifest.parqueteval_manifest.parquet
Token representation
token_store/unit_tokens.uint32.bin
contains concatenated token IDs for retained logical units.
token_store/unit_token_index_v02.parquet
maps each logical corpus unit to its token offset and length and to the corrected v0.2 split.
The token store was created once and reused when the v0.2 split was repaired; the corpus was not retokenized.
Packed data
packed/
contains context-2048 packed binary streams for:
- full train
- full primary eval
- source-family eval slices
- genre eval slices
- author-OOD diagnostic eval
See packed/pack_index.json for exact paths and block counts.
Era metadata
The historical/modern experimental axis is intentionally disabled in this artifact.
The canonical year field was not sufficiently reliable as a representation of the historical era of the underlying literary text, so no historical-vs-modern claim should be derived from it.
Reproducibility
Start with:
DATA_PREP_HANDOFF_v02.json
It records the canonical identity, split policy, tokenizer identity, data-prep identity, token-store paths, packing configuration, and next-phase handoff.
SHA256SUMS.json contains SHA-256 digests and byte sizes for all published artifact files.
Rights / provenance
This repository stores derived tokenized artifacts and metadata from the canonical corpus.
For source-level provenance, rights evidence, licenses, and text metadata, consult the canonical repository:
oliveirabruno01/ptbr-creative-cpt
