datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stories_by_complexityA JSON formatted dataset comprising 31156 short stories for children aged 4 to 16.
This dataset is synthetic and is created using GPT4.
In addition to a story title, summary, and story text, I have included elements such genre, voicing, tense of the story as well as extra information about the age-appropriateness of the story themes (as determined by GPT4), and several text complexity metrics.
The text complexity metrics (Flesch-Kincaid, Gunning Fog Index, SMOG Index, Automated Readability… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/stories_by_complexity.llm-query-complexity-benchmark
LLM Query Complexity Benchmark
A multi-domain, perfectly balanced dataset of 6,000 labeled queries (4,800 train / 1,200 test) for training and evaluating LLM query complexity classifiers that route queries to the most cost-effective inference tier.
Built for the STREAM project (Smart Tiered Routing Engine for AI Models), which routes queries automatically between local CPU models, institutional HPC GPU clusters, and cloud API tiers.
Dataset Summary
Split
Queries… See the full description on the dataset page: https://huggingface.co/datasets/anasnassar/llm-query-complexity-benchmark.ComplexityRouter
Prompt Complexity Dataset
Configurations
Config
File
Description
Size
training
training.jsonl
Training + validation data
4,000
test
test.jsonl
Held‑out test set
400
Data Fields
original_message_id string – the message id from OASST2 (if there was one)
prompt: string – the user prompt
level: string – complexity level (0–3)
category: string – general topic area
reason: string – the reason for giving the level (generated)… See the full description on the dataset page: https://huggingface.co/datasets/RowRed/ComplexityRouter.deita-complexity-scorer-data
Dataset Card for Deita Complexity Scorer Training Data
GitHub | Paper
Deita is an open-sourced project designed to facilitate Automatic Data Selection for instruction tuning in Large Language Models (LLMs).
This dataset includes data for training Deita Complexity Scorer.
Model Family: Other models and the dataset are found in the Deita Collection
Performance
Model
Align
Data Size
MT-Bench
AlpacaEval(%)
OpenLLM (Avg.)
Proprietary Models
GPT-4-Turbo… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/deita-complexity-scorer-data.kazakh-lexical-complexity-classes
Kazakh Lexical Complexity Classes
A CEFR-graded lexical resource for the Kazakh language. The lexicon contains 4,561 lemma–POS entries graded across five CEFR proficiency levels.
Data Format
The dataset is provided as a single JSON file. Each entry has the following fields:
Field
Type
Description
lemma
string
Kazakh word (Cyrillic script)
pos
string
Part of speech (NOUN, VERB, ADJ, ADV, NUM, PRON, OTHER, etc.)
cefr
string
CEFR proficiency level (A1, A2, B1… See the full description on the dataset page: https://huggingface.co/datasets/Gulnur7/kazakh-lexical-complexity-classes.repro-sample-complexity-bounds-for-robust-mean-estimation-with-mean-shift-contaminatio-traces
Agent traces
Agent sessions published from a Trackio Logbook.
swedish-text-complexity
Swedish Text Complexity Dataset
A corpus of Swedish texts annotated with readability and linguistic complexity metrics, created by the Department of Linguistics and Philology at Uppsala University.
Dataset Description
This dataset contains Swedish text passages annotated with multiple complexity metrics, designed to support research in:
Controllable text generation - Train LLMs to generate text at specific reading levels
Educational NLP - Match texts to student reading… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-text-complexity.nepali-Psychology-domain-behaviour-diversity-complexity-sft-dataset
🧠 Nepali Psychology Question Dataset — 2,000 Samples
📌 Overview
The Nepali Psychology Question Dataset is a specialized Nepali-language dataset containing 2,000 psychology-related question-answer records designed for Natural Language Processing (NLP), Large Language Models (LLMs), Small Language Models (SLMs), Supervised Fine-Tuning (SFT), Question Answering (QA), instruction tuning, educational AI, and psychology-domain research.
The dataset is designed with a… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/nepali-Psychology-domain-behaviour-diversity-complexity-sft-dataset.complexity-kink-research
Complexity Kink Research: LLM Code Generation Benchmark
First published: February 22, 2026Author: Michael Hernandez (XxCotHGxX)GitHub: XxCotHGxX/ComplexityKinkLicense: CC BY 4.0
Overview
This dataset supports the Complexity Kink research program — an econometric investigation into whether large language models exhibit a structural performance discontinuity as a function of problem complexity.
The central hypothesis is that LLM code generation performance does not… See the full description on the dataset page: https://huggingface.co/datasets/XxCotHGxX/complexity-kink-research.word_count_word_complexityhan-domestic-task-complexity-annotations-v1
Domestic Task Complexity Annotations
This dataset annotates household tasks
based on their execution complexity
for humanoid robotic systems.
Complexity Dimensions
Physical effort
Precision requirement
Environmental awareness
Use Cases
Task difficulty analysis
Skill learning research
Assistive robotics
Part of
Humanoid Network (HAN)
License
MIT
p2-etf-compression-complexity-resultspt-health-text-complexityPortuguese Health Text Complexity Dataset (PT-PT)
Dataset Summary
The Portuguese Health Text Complexity Dataset (PT-PT) is a curated dataset for text complexity classification in healthcare, focused on European Portuguese.
It combines:
citizen-facing health communication from SNS 24, and
professional clinical language from Direção-Geral da Saúde (DGS),
allowing models to learn the distinction between clear, medium, and complex health-related texts.
Supported Tasks
Text classification
Text… See the full description on the dataset page: https://huggingface.co/datasets/saramscruz/pt-health-text-complexity.
