CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01catherinearnett /apertus_multiblimp0 likes1.4k downloads4mo agoHugging Face02fffoivos /apertus-8b-greek-cpt-modern-greek-train Exact Modern-Greek training content for Apertus 8B Greek CPT This is the public Modern-Greek, train-only document snapshot selected for the full 8B D0 continued-pretraining run. It preserves the upstream v2 schema and metadata; text is reproduced as its exact training-time Apertus-parity PII-masked value. Selection is reconstructed from immutable post-mask training catalogs and content hashes. It contains no replay payload. Exact selected content HPLT Modern… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/apertus-8b-greek-cpt-modern-greek-train.tabular10M<n<100M0 likes342 downloads1mo agoHugging Face03swiss-ai /apertus-pretrain-romanshThis dataset consist of three differnt parts. Monolingual Romansh Data, Polylingual data or more precisely translated data from Romansh into either German, French, Italian or English and Sythetic Data. The Polylingual data consists of aligned and non aligned data. The synthetic data was created by interweaving the translational data and prefacing it with the sentence " This is a text translated from SOURCE LANGUAGE to Rumantsch Grischun". The data has a metadata "idiom" if the if specific… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-romansh.tabulartranslation100K<n<1M4 likes216 downloads1y agoHugging Face04swiss-ai /apertus-pretrain-swiss Swiss Pretrain Data This dataset provides a large collection of open-access and license-compliant Swiss data sources for language model training. The dataset includes the following sources: Name Internal ID Tokens (B) Description Curia Vista curiavista 0.5 Legal and administrative documents from the Swiss database of parliamentary proceedings. enscheidsuche enscheidsuche_html 4.5 Swiss court decisions, sampled at 50% for balance. FineWeb-2… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-swiss.textfill-mask1M<n<10M8 likes193 downloads1y agoHugging Face05swiss-ai /Apertus_v1.5_Preference_Data Apertus 1.5 Preference Dataset This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model. The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us. How this dataset was built Prompts. Taken from Dolci-Instruct-DPO (ODC-BY). Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus_v1.5_Preference_Data.tabulartext-generation100K<n<1M4 likes165 downloads29d agoHugging Face06swiss-ai /apertus-sft-mixture Apertus Supervised Finetuning Data Our supervised finetuning data contains a carefully curated blend of instruction-following datasets, developed through eight iterations of empirical evaluation. This final mixture comprises approximately 3.8 million examples from diverse sources, balancing generalinstruction-following, mathematical reasoning, code generation, and multilingual capabilities. More details about data provenance, preparation, and statistics can be found in our tech… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-sft-mixture.texttext-generation1M<n<10M7 likes157 downloads1y agoHugging Face07juiceb0xc0de /apertus-v1.1-1.5b-atlas0 likes156 downloads3mo agoHugging Face08daslab-testing /Apertus-8B-2509-microQAT-logitsThis dataset provides a small sample of TOP-K logits computed using swiss-ai/Apertus-8B-2509 on samples from Data Phase 5 of Apertus pre-training. Format This data represents documents packed into chuncks of 4096 tokens separated by EOS. The provided fields are as follows: input_ids: Input tokens. index: Positions of top-256 highest-probability next-token predictions for each token. exp_logits: Normalized probabilities of top-256 highest-probability next-token predictions for each… See the full description on the dataset page: https://huggingface.co/datasets/daslab-testing/Apertus-8B-2509-microQAT-logits.text-generation10K<n<100K0 likes125 downloads6mo agoHugging Face09mlx-community /Apertus-v1.5-QAT-10K mlx-community/Apertus-v1.5-QAT-10K This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models. texttext-generation10K<n<100K1 likes112 downloads5d agoHugging Face10swiss-ai /apertus-pretrain-poisonandcanariesThis dataset was used as part of Apertus v1 training for poisoning experiments. See our technical report for details, as well as the dedicated study. tabular100K<n<1M3 likes88 downloads11mo agoHugging Face11juiceb0xc0de /apertus-v1.1-0.5b-atlas apertus-v1.1-0.5b-atlas image100K<n<1M0 likes78 downloads25d agoHugging Face12thomaskiefer /EAGLE3-Apertus-8B-Instruct-2509-Data EAGLE3-Apertus-8B-Instruct-2509-Data Training dataset for the thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509 speculative decoding draft model. Dataset Description This dataset contains ~375k multi-turn conversations used to train an Eagle3 draft model for swiss-ai/Apertus-8B-Instruct-2509. Data Sources The prompts are sourced from: UltraChat - Large-scale multi-turn dialogue dataset ShareGPT - Real user conversations OpenThoughts-114k-math - Mathematical… See the full description on the dataset page: https://huggingface.co/datasets/thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509-Data.texttext-generation100K<n<1M0 likes53 downloads10mo agoHugging Face13swiss-ai /apertus-posttrain-romansh license: cc-by-4.0 Romansh SFT Data Supervised fine-tuning (SFT) splits built from the swiss-ai/apertus-pretrain-rumansh corpus. It contains dictionary list translation, sentence-level translation, idiom identification, and a small set of human-translated Romansh instructions. Source hub: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-rumansh Provenance Dictionaries: All dictionary entries originate from Pledarigrond and are provided by the… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-posttrain-romansh.text10K<n<100K5 likes42 downloads1y agoHugging Face14daslab-testing /Apertus-8B-Instruct-2509-microQAT-logits100K<n<1M0 likes32 downloads5mo agoHugging Face15jvamvas /apertus-pretrain-romansh-backtranslatedVersion of https://hf.co/datasets/swiss-ai/apertus-pretrain-romansh (monolingual split only) that includes MT-generated translations into German. The intended purpose of this dataset is to train MT systems or LLMs on the task of idiom-specific German→Romansh translation. Note that the German translations in this dataset might contain errors, since they have been automatically generated by an MT system. Composition of the dataset and Romansh data sources See… See the full description on the dataset page: https://huggingface.co/datasets/jvamvas/apertus-pretrain-romansh-backtranslated.tabular100K<n<1M1 likes29 downloads2mo agoHugging Face16JingweiNi /train_prm800k_apertus_70b_gpt-oss-120b_annotated_256tok_part6text1K<n<10K0 likes28 downloads5mo agoHugging Face17fffoivos /greek-apertus-sftgated Greek Apertus SFT datasets The supervised fine-tuning data of the Greek Apertus project: the GlossAPI team of EELLAK (Open Technologies Alliance) continues the pre-training of swiss-ai/Apertus-8B-2509 on Greek text and then trains it on instructions, with a grant from the Swiss AI Initiative. This repository holds every training arm we assembled, exactly as it went (or goes) to the trainer: one messages list per row, chat format, no system turn. Access is gated: request it and… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-apertus-sft.texttext-generation100K<n<1M0 likes28 downloads15d agoHugging Face18juiceb0xc0de /apertus-v1.1-4b-atlas0 likes27 downloads3mo agoHugging Face19liechticonsulting /filtered_apertus_pretrainedUsed to create this dataset: import json import os from datasets import load_dataset from tqdm import tqdm # --- Configuration --- DATASET_NAME = "swiss-ai/apertus-pretrain-swiss" SUBSTRING_TO_FILTER = "entscheidsuche_html" COLUMN_TO_CHECK = "id" OUTPUT_FILENAME = "filtered_apertus_pretrain_swiss.jsonl" def filter_function(example): """ Returns True to keep the example, False to discard it. We keep the row only if the substring is NOT in the 'id' column. """ return… See the full description on the dataset page: https://huggingface.co/datasets/liechticonsulting/filtered_apertus_pretrained.text100K<n<1M0 likes22 downloads1y agoHugging Face20snae /emotion_stories_Apertus_8B_Instruct Emotion Stories — Apertus-8B-Instruct Synthetic short stories that convey a target emotion implicitly — without ever naming the emotion or its direct synonyms. Each story expresses the emotion only through actions, body language, dialogue, internal reactions, and situational context. The dataset was built to study emotion representations in language models (e.g. probing and activation-steering experiments). Generated with swiss-ai/Apertus-8B-Instruct-2509. A companion set… See the full description on the dataset page: https://huggingface.co/datasets/snae/emotion_stories_Apertus_8B_Instruct.texttext-generationn<1K0 likes20 downloads3mo agoHugging Face21danish-foundation-models /croco-munin-apertus-8b-da-simpo-fulltabular1K<n<10K0 likes19 downloads2mo agoHugging Face22danish-foundation-models /croco-munin-apertus-8b-da-50ktabular10K<n<100K0 likes19 downloads2mo agoHugging Face23danish-foundation-models /croco-munin-apertus-8b-da-simpo-full-50ktabular10K<n<100K0 likes18 downloads2mo agoHugging Face24jvamvas /apertus-pretrain-romansh-backtranslated-sftSubsampled version of https://huggingface.co/datasets/jvamvas/apertus-pretrain-romansh-backtranslated that has been processed as follows: Balanced subsampling to 30k samples (5k samples per variety) Samples with higher LID scores are prioritized Formatted as German-to-Romansh translation instruction pairs in prompt/completion format (prompt: Übersetze den folgenden Text nach {variety}:\n\n{german_backtranslation}; completion: the Romansh text) Normalized linebreaks to have clear text… See the full description on the dataset page: https://huggingface.co/datasets/jvamvas/apertus-pretrain-romansh-backtranslated-sft.text10K<n<100K0 likes17 downloads6mo agoHugging Face25JingweiNi /train_prm800k_apertus_70b_math_texts_part1text1K<n<10K0 likes17 downloads5mo agoHugging Face26JingweiNi /apertus_70b_3899_0.2_0.75_correctness_no_refuse_sfttabular10K<n<100K0 likes15 downloads1y agoHugging Face27JingweiNi /train_prm800k_apertus_70b_math_texts_part5text1K<n<10K0 likes15 downloads5mo agoHugging Face28JingweiNi /train_prm800k_apertus_70b_math_texts_256tok_part6text1K<n<10K0 likes15 downloads5mo agoHugging Face29JingweiNi /train_prm800k_apertus_70b_gpt-oss-120b_annotated_256tok_part4text1K<n<10K0 likes15 downloads5mo agoHugging Face30JingweiNi /train_prm800k_apertus_70b_gpt-oss-120b_annotated_256tok_part7text1K<n<10K0 likes15 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.