datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swedish-text-complexity
Swedish Text Complexity Dataset
A corpus of Swedish texts annotated with readability and linguistic complexity metrics, created by the Department of Linguistics and Philology at Uppsala University.
Dataset Description
This dataset contains Swedish text passages annotated with multiple complexity metrics, designed to support research in:
Controllable text generation - Train LLMs to generate text at specific reading levels
Educational NLP - Match texts to student reading… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-text-complexity.pt-health-text-complexityPortuguese Health Text Complexity Dataset (PT-PT)
Dataset Summary
The Portuguese Health Text Complexity Dataset (PT-PT) is a curated dataset for text complexity classification in healthcare, focused on European Portuguese.
It combines:
citizen-facing health communication from SNS 24, and
professional clinical language from Direção-Geral da Saúde (DGS),
allowing models to learn the distinction between clear, medium, and complex health-related texts.
Supported Tasks
Text classification
Text… See the full description on the dataset page: https://huggingface.co/datasets/saramscruz/pt-health-text-complexity.
