datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lls-complexitybreakeven_complexityThis repository contains the dataset for "Breakeven complexity: A new perspective on neural partial differential equation solvers".
Dataset detail:
Navier-Stokes: Simulated via Exponax. Contains field "u". 20,000 training trajectories and 1,000 test trajectories. Shape (N, res, res, T).
Kuramoto-Sivashinsky: Simulated via Exponax. Contains field "u". 20,000 training trajectories and 1,000 test trajectories. Shape (N, res, res, T).
Gray-Scott: Simulated via Exponax. Contains field "u" and "v".… See the full description on the dataset page: https://huggingface.co/datasets/yijingz/breakeven_complexity.complexity-atlas-posttrain
Complexity Atlas Posttrain — Card Corpus V2
An English supervised fine-tuning corpus generated from authored semantic
frames, role-separated prompt/answer/thinking plans, compatibility graphs, and
VariableBy2D reservoirs. All 15 task families, including natural dialogue,
belong to one audited corpus and one tokenizer-compatible training view.
Release
Split
Examples
Train
224,654
Validation
2,478
Test
1,894
Total
229,026
The generator renders… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/complexity-atlas-posttrain.question-type-and-complexity
Question Type and Complexity (QTC) Dataset
Dataset Overview
The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features.
Key Features:
2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.Orion-Creative_Writing-Complexitystories_by_complexityA JSON formatted dataset comprising 31156 short stories for children aged 4 to 16.
This dataset is synthetic and is created using GPT4.
In addition to a story title, summary, and story text, I have included elements such genre, voicing, tense of the story as well as extra information about the age-appropriateness of the story themes (as determined by GPT4), and several text complexity metrics.
The text complexity metrics (Flesch-Kincaid, Gunning Fog Index, SMOG Index, Automated Readability… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/stories_by_complexity.the-complexity-trap
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
This dataset contains our raw experimental data (ie. agent trajectories) accompanying the paper "The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management" and Tobias Lindenbauer's Master's thesis.
The data in this repository are compressed to .tar.gz archives. For detailed instructions on how to use these data… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/the-complexity-trap.ComplexityMT
ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation
Official data release for the paper "ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation" (Imperial et al., 2026). ComplexityMT is a benchmark for studying how text complexity — operationalised through the Common European Framework of Reference (CEFR) — interacts with machine translation across six languages and five MT systems.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/ComplexityMT.faitheval-counterfactual-complexity-3-llama70b-balancedllm-query-complexity-benchmark
LLM Query Complexity Benchmark
A multi-domain, perfectly balanced dataset of 6,000 labeled queries (4,800 train / 1,200 test) for training and evaluating LLM query complexity classifiers that route queries to the most cost-effective inference tier.
Built for the STREAM project (Smart Tiered Routing Engine for AI Models), which routes queries automatically between local CPU models, institutional HPC GPU clusters, and cloud API tiers.
Dataset Summary
Split
Queries… See the full description on the dataset page: https://huggingface.co/datasets/anasnassar/llm-query-complexity-benchmark.swedish-cefr-text-complexity
Swedish CEFR Text Complexity Dataset
This dataset contains Swedish text examples labeled with approximate CEFR
reading levels from A1 to C2.
It was created for an information retrieval assignment about training text
classifiers with embeddings. The companion demo and classifier use
nicher92/saga-embed_v1 sentence embeddings and classical scikit-learn
classifiers.
The dataset is intended for Swedish text-complexity classification: given a
short Swedish sentence or paragraph, predict… See the full description on the dataset page: https://huggingface.co/datasets/kvest/swedish-cefr-text-complexity.quantum-information-and-complexity-theory
Neura Parse — Quantum Information & Complexity Theory: Channels, Entropies, Classes & the Structure of Advantage
A proof-based theoretical-foundations vertical uniting quantum information theory (channels, entropies, entanglement measures, distinguishability, capacities, Shannon theory) with quantum complexity theory and the structure of quantum advantage (classes, Hamiltonian complexity, sampling-based advantage and its verification, pseudorandomness, dequantization).… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-information-and-complexity-theory.Meta-Llama-3-8B-Instruct_ultrafeedback-annotate-judge-mtbench_cot_helpsteer_complexitystl_high_complexitybrick-complexity-extractor
🧱 Brick Complexity Extractor Dataset
76,831 user queries labeled by complexity for LLM routing
Regolo.ai · Model · Brick SR1 on GitHub · API Docs
Overview
This dataset provides 76,831 user queries annotated with a complexity label (easy, medium, or hard) indicating the cognitive effort and reasoning depth required to answer each query. It was created to train the Brick Complexity Extractor, a LoRA adapter used in the Brick Semantic Router for… See the full description on the dataset page: https://huggingface.co/datasets/regolo/brick-complexity-extractor.ComplexityRouter
Prompt Complexity Dataset
Configurations
Config
File
Description
Size
training
training.jsonl
Training + validation data
4,000
test
test.jsonl
Held‑out test set
400
Data Fields
original_message_id string – the message id from OASST2 (if there was one)
prompt: string – the user prompt
level: string – complexity level (0–3)
category: string – general topic area
reason: string – the reason for giving the level (generated)… See the full description on the dataset page: https://huggingface.co/datasets/RowRed/ComplexityRouter.deita-complexity-scorer-data
Dataset Card for Deita Complexity Scorer Training Data
GitHub | Paper
Deita is an open-sourced project designed to facilitate Automatic Data Selection for instruction tuning in Large Language Models (LLMs).
This dataset includes data for training Deita Complexity Scorer.
Model Family: Other models and the dataset are found in the Deita Collection
Performance
Model
Align
Data Size
MT-Bench
AlpacaEval(%)
OpenLLM (Avg.)
Proprietary Models
GPT-4-Turbo… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/deita-complexity-scorer-data.complexity-messagestruthfulqa-complexity-3-llama70b-balancedcomplexity-atlas-images
Complexity Atlas Images
Complexity Atlas Images is a provenance-preserving CC0 image-text bank for
image captioning and multimodal research. Version 1.0.0 contains 336,245
normalized museum images and English metadata-derived captions.
Splits
Split
Samples
Train
329,510
Validation
3,428
Test
3,307
Total
336,245
The dataset occupies approximately 6.67 GB across 68 WebDataset TAR shards.
Source
This release is derived from… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/complexity-atlas-images.kazakh-lexical-complexity-classes
Kazakh Lexical Complexity Classes
A CEFR-graded lexical resource for the Kazakh language. The lexicon contains 4,561 lemma–POS entries graded across five CEFR proficiency levels.
Data Format
The dataset is provided as a single JSON file. Each entry has the following fields:
Field
Type
Description
lemma
string
Kazakh word (Cyrillic script)
pos
string
Part of speech (NOUN, VERB, ADJ, ADV, NUM, PRON, OTHER, etc.)
cefr
string
CEFR proficiency level (A1, A2, B1… See the full description on the dataset page: https://huggingface.co/datasets/Gulnur7/kazakh-lexical-complexity-classes.repro-sample-complexity-bounds-for-robust-mean-estimation-with-mean-shift-contaminatio-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-the-optimal-sample-complexity-of-linear-contracts-bundle
Reproduction: The Optimal Sample Complexity of Linear Contracts
js-function-complexityclaude-code-project-complexity-workflowscomplexity-atlas-image-edits
Complexity Atlas Image Edits
Complexity Atlas Image Edits is an aligned instruction-guided image editing
dataset derived from the normalized public-domain Complexity Atlas image bank.
It contains 336,245 explicit source/instruction/target triplets at 256 x 256.
Dataset structure
Each WebDataset record contains:
<edit_id>.source.webp
<edit_id>.target.webp
<edit_id>.txt
<edit_id>.json
source.webp is a deterministic degraded image;
target.webp is the unchanged… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/complexity-atlas-image-edits.databricks-dolly15k-semantic-complexity
Databricks - Dolly 15k – Enriched Variant (Instruction-Tuned with Semantic and Complexity Augmentation)
Overview
This dataset is a semantically enriched and complexity-aware extension of the original Databricks Dolly 15k, purpose-built for evaluating and training instruction-following models. Each sample is augmented with additional signals to enable more nuanced filtering, curriculum learning, and benchmark development across diverse NLP tasks.
Dataset Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/databricks-dolly15k-semantic-complexity.swedish-text-complexity
Swedish Text Complexity Dataset
A corpus of Swedish texts annotated with readability and linguistic complexity metrics, created by the Department of Linguistics and Philology at Uppsala University.
Dataset Description
This dataset contains Swedish text passages annotated with multiple complexity metrics, designed to support research in:
Controllable text generation - Train LLMs to generate text at specific reading levels
Educational NLP - Match texts to student reading… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-text-complexity.LlamaGemma-GSM8K-MMLU-Preference-16K-Eval-Complexityself_instruct_with_complexity_scores
