datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swedish-medical-exams-mcq-1006-json
Dataset Card for Swedish Medical Exam MCQs
Dataset Description
This dataset contains multiple-choice questions from Swedish medical exams.
Languages
The dataset is in Swedish (sv).
Dataset Structure
Each entry in the dataset contains the following fields:
question: The question
options: An array of possible answers
answer: The correct answer
language: The language of the question (always "sv" for Swedish)
country: The country of origin (always… See the full description on the dataset page: https://huggingface.co/datasets/sarafuyu/swedish-medical-exams-mcq-1006-json.swedish-construction-faq
Swedish Construction FAQ — Open Q&A Dataset
Open bilingual (Swedish/English) Q&A dataset for the Swedish construction
industry (byggbranschen). 503 Q&A pairs across 39 categories, every answer
grounded in Swedish primary law and authoritative guidance.
Maintained by Zaragoza AB, Helsingborg, Sweden.
DOI: 10.5281/zenodo.19630803 ·
Wikidata: Q139393633
Try it first
Live search demo: huggingface.co/spaces/DecDEPO/swedish-construction-faq-search
Colab quickstart: Open in… See the full description on the dataset page: https://huggingface.co/datasets/DecDEPO/swedish-construction-faq.Alpaca-Lora-GPT4-Swedish-RefinedThis is based on: https://huggingface.co/datasets/jeremyc/Alpaca-Lora-GPT4-Swedish
I've done extensive cleaning (but I'm not yet done).
This includes:
Purging erroneous and sometimes offensive generations by the translator
Fixing code instances up to row 10300. All code was botched. There may still be some html instances to fix, but at least all python should be valid.
uner_llm_inst_swedish
Dataset Card for Universal NER v1 in the Aya format - Swedish subset
This dataset is a format conversion for the Swedish data in the original Universal NER v1 into the Aya instruction format and it's released here under the same CC-BY-SA 4.0 license and conditions.
The dataset contains different subsets and their dev/test/train splits, depending on language. For more details, please refer to:
Dataset Details
For the original Universal NER dataset v1 and more details… See the full description on the dataset page: https://huggingface.co/datasets/universalner/uner_llm_inst_swedish.general-qa-swedishgovernment-interpellation-qa-swedishopen_assistant_swedishclassical-swedish-allusions-gold
Classical-Swedish Allusions: Gold Benchmark
A small curated reference set of cross-lingual allusions from classical Greek and Latin into Swedish literature. Used as a held-out evaluation set for the Ericu950/classical-swedish-citations-v2 project.
Purpose
This is not training data. It is a fixed, public reference standard. Our cross-lingual allusion-detection system will aspire to recover these. The set was archived publicly before evaluating any system against it.
The… See the full description on the dataset page: https://huggingface.co/datasets/Ericu950/classical-swedish-allusions-gold.swedish_healthcare_dpo_sft_dataset
Swedish Healthcare DPO/SFT Dataset
This dataset contains Swedish conversational prompts paired with chosen (preferred) and rejected (dispreferred) responses, along with English translations. It is designed for fine-tuning Large Language Models (LLMs) using preference-based methods. The (prompt, chosen, rejected) structure makes it particularly well-suited for Direct Preference Optimization (DPO) and similar algorithms.
Sweden is renowned for its high standards in healthcare and… See the full description on the dataset page: https://huggingface.co/datasets/rzgar/swedish_healthcare_dpo_sft_dataset.swedish-health-source-triage
Swedish Health Source Triage
This is a small custom text-classification dataset for an Information Retrieval
assignment about embeddings. The task is to classify short health-information
texts by the public source family they resemble:
1177.se: patient-facing healthcare guidance
socialstyrelsen.se: national authority reports, guidelines, and statistics
lakemedelsverket.se: medicine and medical-product regulation
Intended use
The dataset is intentionally compact and… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/swedish-health-source-triage.swedish-pre-a1-learning-agent-dataset
Dataset Card: Swedish Beginner Agent Structured Dataset
Dataset Summary
This dataset is a structured Pre-A1 Swedish learning resource for a small learning agent. It includes vocabulary, sentence examples, dialogue function definitions, and dialogue planning sequences across six daily-life scenarios.
Languages
Swedish (sv)
English translations (en)
Chinese translations (zh)
Dataset Structure
vocabulary.jsonl: vocabulary items… See the full description on the dataset page: https://huggingface.co/datasets/SeanSha30/swedish-pre-a1-learning-agent-dataset.swedish-text-complexity
Swedish Text Complexity Dataset
A corpus of Swedish texts annotated with readability and linguistic complexity metrics, created by the Department of Linguistics and Philology at Uppsala University.
Dataset Description
This dataset contains Swedish text passages annotated with multiple complexity metrics, designed to support research in:
Controllable text generation - Train LLMs to generate text at specific reading levels
Educational NLP - Match texts to student reading… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-text-complexity.
