datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apertus-pretrain-romanshThis dataset consist of three differnt parts. Monolingual Romansh Data, Polylingual data or more precisely translated data from Romansh into either German, French, Italian or English and Sythetic Data.
The Polylingual data consists of aligned and non aligned data. The synthetic data was created by interweaving the translational data and prefacing it with the sentence " This is a text translated from SOURCE LANGUAGE to Rumantsch Grischun".
The data has a metadata "idiom" if the if specific… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-romansh.Apertus-v1.5-QAT-10K
mlx-community/Apertus-v1.5-QAT-10K
This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models.
apertus-pretrain-poisonandcanariesThis dataset was used as part of Apertus v1 training for poisoning experiments. See our technical report for details, as well as the dedicated study.
EAGLE3-Apertus-8B-Instruct-2509-Data
EAGLE3-Apertus-8B-Instruct-2509-Data
Training dataset for the thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509 speculative decoding draft model.
Dataset Description
This dataset contains ~375k multi-turn conversations used to train an Eagle3 draft model for swiss-ai/Apertus-8B-Instruct-2509.
Data Sources
The prompts are sourced from:
UltraChat - Large-scale multi-turn dialogue dataset
ShareGPT - Real user conversations
OpenThoughts-114k-math - Mathematical… See the full description on the dataset page: https://huggingface.co/datasets/thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509-Data.apertus-posttrain-romansh
license: cc-by-4.0
Romansh SFT Data
Supervised fine-tuning (SFT) splits built from the swiss-ai/apertus-pretrain-rumansh corpus. It contains dictionary list translation, sentence-level translation, idiom identification, and a small set of human-translated Romansh instructions.
Source hub: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-rumansh
Provenance
Dictionaries: All dictionary entries originate from Pledarigrond and are provided by the… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-posttrain-romansh.apertus-pretrain-romansh-backtranslatedVersion of https://hf.co/datasets/swiss-ai/apertus-pretrain-romansh (monolingual split only) that includes MT-generated translations into German.
The intended purpose of this dataset is to train MT systems or LLMs on the task of idiom-specific German→Romansh translation. Note that the German translations in this dataset might contain errors, since they have been automatically generated by an MT system.
Composition of the dataset and Romansh data sources
See… See the full description on the dataset page: https://huggingface.co/datasets/jvamvas/apertus-pretrain-romansh-backtranslated.greek-apertus-sft
Greek Apertus SFT datasets
The supervised fine-tuning data of the Greek Apertus project: the GlossAPI team of EELLAK (Open Technologies Alliance) continues the pre-training of swiss-ai/Apertus-8B-2509 on Greek text and then trains it on instructions, with a grant from the Swiss AI Initiative. This repository holds every training arm we assembled, exactly as it went (or goes) to the trainer: one messages list per row, chat format, no system turn.
Access is gated: request it and… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-apertus-sft.filtered_apertus_pretrainedUsed to create this dataset:
import json
import os
from datasets import load_dataset
from tqdm import tqdm
# --- Configuration ---
DATASET_NAME = "swiss-ai/apertus-pretrain-swiss"
SUBSTRING_TO_FILTER = "entscheidsuche_html"
COLUMN_TO_CHECK = "id"
OUTPUT_FILENAME = "filtered_apertus_pretrain_swiss.jsonl"
def filter_function(example):
"""
Returns True to keep the example, False to discard it.
We keep the row only if the substring is NOT in the 'id' column.
"""
return… See the full description on the dataset page: https://huggingface.co/datasets/liechticonsulting/filtered_apertus_pretrained.emotion_stories_Apertus_8B_Instruct
Emotion Stories — Apertus-8B-Instruct
Synthetic short stories that convey a target emotion implicitly — without ever
naming the emotion or its direct synonyms. Each story expresses the emotion only
through actions, body language, dialogue, internal reactions, and situational
context. The dataset was built to study emotion representations in language
models (e.g. probing and activation-steering experiments).
Generated with swiss-ai/Apertus-8B-Instruct-2509.
A companion set… See the full description on the dataset page: https://huggingface.co/datasets/snae/emotion_stories_Apertus_8B_Instruct.croco-munin-apertus-8b-da-simpo-fullcroco-munin-apertus-8b-da-50kcroco-munin-apertus-8b-da-simpo-full-50kapertus-pretrain-romansh-backtranslated-sftSubsampled version of https://huggingface.co/datasets/jvamvas/apertus-pretrain-romansh-backtranslated that has been processed as follows:
Balanced subsampling to 30k samples (5k samples per variety)
Samples with higher LID scores are prioritized
Formatted as German-to-Romansh translation instruction pairs in prompt/completion format (prompt: Übersetze den folgenden Text nach {variety}:\n\n{german_backtranslation}; completion: the Romansh text)
Normalized linebreaks to have clear text… See the full description on the dataset page: https://huggingface.co/datasets/jvamvas/apertus-pretrain-romansh-backtranslated-sft.croco-munin-apertus-8b-da-generatedcroco-munin-apertus-8b-da-goldapertus-c3-dedup-audit-dedup-20260519t010924z
Apertus C3 Dedup Audit - dedup_20260519T010924Z
This gated public dataset repository contains machine-readable duplicate-overlap artifacts for the run dedup_20260519T010924Z.
Scope: HF source-pool overlap with Apertus pretraining sources, not the exact sampled C3 training mix. The held-out contamination check was skipped because no holdout doc-key list was provided.
Headline Counts
HF source-pool docs audited: 98,203,721
Matched HF source-pool docs: 2,223,781… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/apertus-c3-dedup-audit-dedup-20260519t010924z.croco-munin-apertus-8b-da-simpo-tunedapertus-annotation-feedbackcroco-munin-apertus-8b-dacroco-munin-apertus-8b-da-simpocroco-munin-apertus-8b-da-lsapertus-rollouts-1
