fffoivos/apertus-8b-greek-cpt-modern-greek-train
Exact Modern-Greek training content for Apertus 8B Greek CPT This is the public Modern-Greek, train-only document snapshot selected for the full 8B D0 continued-pretraining run. It preserves the upstream v2 schema and metadata; text is reproduced as its exact training-time Apertus-parity PII-masked value. Selection is reconstructed from immutable post-mask training catalogs and content hashes. It contains no replay payload. Exact selected content HPLT Modern… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/apertus-8b-greek-cpt-modern-greek-train.
Exact Modern-Greek training content for Apertus 8B Greek CPT
This is the public Modern-Greek, train-only document snapshot selected for the full 8B D0 continued-pretraining run. It preserves the upstream v2 schema and metadata; text is reproduced as its exact training-time Apertus-parity PII-masked value. Selection is reconstructed from immutable post-mask training catalogs and content hashes. It contains no replay payload.
Exact selected content
- HPLT Modern Greek: 46,535,439 documents; 41,512,804,679 active training tokens.
- GlossAPI/non-HPLT Modern Greek: 3,086,052 documents; 19,068,732,797 active training tokens.
- Total: 49,621,491 documents; 60,581,537,476 active training tokens.
Processing and provenance
The upstream corpus was revision-pinned. The training workflow applied heldout exclusion, GreekMMLU decontamination, Apertus-standard email/IP/validated-IBAN masking, and global exact post-mask deduplication before tokenization. This release reapplies that frozen masking function only to reproduce the selected training text; it does not introduce another policy, deduplication, or retokenization pass.
The companion private dataset fffoivos/apertus-8b-greek-cpt-d0-full-mix contains the complete packed 79/20/1 mixture and its restricted replay provenance. This public dataset does not grant redistribution rights for those replay sources.
