text complexity
swedish-cefr-text-complexity
Swedish CEFR Text Complexity Dataset
This dataset contains Swedish text examples labeled with approximate CEFR
reading levels from A1 to C2.
It was created for an information retrieval assignment about training text
classifiers with embeddings. The companion demo and classifier use
nicher92/saga-embed_v1 sentence embeddings and classical scikit-learn
classifiers.
The dataset is intended for Swedish text-complexity classification: given a
short Swedish sentence or paragraph, predict… See the full description on the dataset page: https://huggingface.co/datasets/kvest/swedish-cefr-text-complexity.swedish-text-complexity
Swedish Text Complexity Dataset
A corpus of Swedish texts annotated with readability and linguistic complexity metrics, created by the Department of Linguistics and Philology at Uppsala University.
Dataset Description
This dataset contains Swedish text passages annotated with multiple complexity metrics, designed to support research in:
Controllable text generation - Train LLMs to generate text at specific reading levels
Educational NLP - Match texts to student reading… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-text-complexity.parallel-complexity-med-textpt-health-text-complexityPortuguese Health Text Complexity Dataset (PT-PT)
Dataset Summary
The Portuguese Health Text Complexity Dataset (PT-PT) is a curated dataset for text complexity classification in healthcare, focused on European Portuguese.
It combines:
citizen-facing health communication from SNS 24, and
professional clinical language from Direção-Geral da Saúde (DGS),
allowing models to learn the distinction between clear, medium, and complex health-related texts.
Supported Tasks
Text classification
Text… See the full description on the dataset page: https://huggingface.co/datasets/saramscruz/pt-health-text-complexity.
