datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apertus-pretrain-swiss
Swiss Pretrain Data
This dataset provides a large collection of open-access and license-compliant Swiss data sources for language model training.
The dataset includes the following sources:
Name
Internal ID
Tokens (B)
Description
Curia Vista
curiavista
0.5
Legal and administrative documents from the Swiss database of parliamentary proceedings.
enscheidsuche
enscheidsuche_html
4.5
Swiss court decisions, sampled at 50% for balance.
FineWeb-2… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-swiss.apertus-sft-mixture
Apertus Supervised Finetuning Data
Our supervised finetuning data contains a carefully curated blend of instruction-following datasets,
developed through eight iterations of empirical evaluation. This final mixture comprises approximately
3.8 million examples from diverse sources, balancing generalinstruction-following, mathematical reasoning,
code generation, and multilingual capabilities.
More details about data provenance, preparation, and statistics can be found in our tech… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-sft-mixture.Apertus_v1.5_Preference_Data
Apertus 1.5 Preference Dataset
This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model.
The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us.
How this dataset was built
Prompts. Taken from Dolci-Instruct-DPO (ODC-BY).
Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus_v1.5_Preference_Data.Apertus-v1.5-QAT-10K
mlx-community/Apertus-v1.5-QAT-10K
This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models.
Apertus-8B-2509-microQAT-logitsThis dataset provides a small sample of TOP-K logits computed using swiss-ai/Apertus-8B-2509 on samples from Data Phase 5 of Apertus pre-training.
Format
This data represents documents packed into chuncks of 4096 tokens separated by EOS. The provided fields are as follows:
input_ids: Input tokens.
index: Positions of top-256 highest-probability next-token predictions for each token.
exp_logits: Normalized probabilities of top-256 highest-probability next-token predictions for each… See the full description on the dataset page: https://huggingface.co/datasets/daslab-testing/Apertus-8B-2509-microQAT-logits.EAGLE3-Apertus-8B-Instruct-2509-Data
EAGLE3-Apertus-8B-Instruct-2509-Data
Training dataset for the thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509 speculative decoding draft model.
Dataset Description
This dataset contains ~375k multi-turn conversations used to train an Eagle3 draft model for swiss-ai/Apertus-8B-Instruct-2509.
Data Sources
The prompts are sourced from:
UltraChat - Large-scale multi-turn dialogue dataset
ShareGPT - Real user conversations
OpenThoughts-114k-math - Mathematical… See the full description on the dataset page: https://huggingface.co/datasets/thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509-Data.greek-apertus-sft
Greek Apertus SFT datasets
The supervised fine-tuning data of the Greek Apertus project: the GlossAPI team of EELLAK (Open Technologies Alliance) continues the pre-training of swiss-ai/Apertus-8B-2509 on Greek text and then trains it on instructions, with a grant from the Swiss AI Initiative. This repository holds every training arm we assembled, exactly as it went (or goes) to the trainer: one messages list per row, chat format, no system turn.
Access is gated: request it and… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-apertus-sft.emotion_stories_Apertus_8B_Instruct
Emotion Stories — Apertus-8B-Instruct
Synthetic short stories that convey a target emotion implicitly — without ever
naming the emotion or its direct synonyms. Each story expresses the emotion only
through actions, body language, dialogue, internal reactions, and situational
context. The dataset was built to study emotion representations in language
models (e.g. probing and activation-steering experiments).
Generated with swiss-ai/Apertus-8B-Instruct-2509.
A companion set… See the full description on the dataset page: https://huggingface.co/datasets/snae/emotion_stories_Apertus_8B_Instruct.
