apertus
Datasets
All datasets matching “apertus”apertus_multiblimpapertus-8b-greek-cpt-modern-greek-train
Exact Modern-Greek training content for Apertus 8B Greek CPT
This is the public Modern-Greek, train-only document snapshot selected for the full 8B D0 continued-pretraining run. It preserves the upstream v2 schema and metadata; text is reproduced as its exact training-time Apertus-parity PII-masked value. Selection is reconstructed from immutable post-mask training catalogs and content hashes. It contains no replay payload.
Exact selected content
HPLT Modern… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/apertus-8b-greek-cpt-modern-greek-train.apertus-pretrain-romanshThis dataset consist of three differnt parts. Monolingual Romansh Data, Polylingual data or more precisely translated data from Romansh into either German, French, Italian or English and Sythetic Data.
The Polylingual data consists of aligned and non aligned data. The synthetic data was created by interweaving the translational data and prefacing it with the sentence " This is a text translated from SOURCE LANGUAGE to Rumantsch Grischun".
The data has a metadata "idiom" if the if specific… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-romansh.apertus-pretrain-swiss
Swiss Pretrain Data
This dataset provides a large collection of open-access and license-compliant Swiss data sources for language model training.
The dataset includes the following sources:
Name
Internal ID
Tokens (B)
Description
Curia Vista
curiavista
0.5
Legal and administrative documents from the Swiss database of parliamentary proceedings.
enscheidsuche
enscheidsuche_html
4.5
Swiss court decisions, sampled at 50% for balance.
FineWeb-2… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-swiss.Apertus_v1.5_Preference_Data
Apertus 1.5 Preference Dataset
This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model.
The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us.
How this dataset was built
Prompts. Taken from Dolci-Instruct-DPO (ODC-BY).
Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus_v1.5_Preference_Data.apertus-sft-mixture
Apertus Supervised Finetuning Data
Our supervised finetuning data contains a carefully curated blend of instruction-following datasets,
developed through eight iterations of empirical evaluation. This final mixture comprises approximately
3.8 million examples from diverse sources, balancing generalinstruction-following, mathematical reasoning,
code generation, and multilingual capabilities.
More details about data provenance, preparation, and statistics can be found in our tech… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-sft-mixture.
