complexity
lls-complexitybreakeven_complexityThis repository contains the dataset for "Breakeven complexity: A new perspective on neural partial differential equation solvers".
Dataset detail:
Navier-Stokes: Simulated via Exponax. Contains field "u". 20,000 training trajectories and 1,000 test trajectories. Shape (N, res, res, T).
Kuramoto-Sivashinsky: Simulated via Exponax. Contains field "u". 20,000 training trajectories and 1,000 test trajectories. Shape (N, res, res, T).
Gray-Scott: Simulated via Exponax. Contains field "u" and "v".… See the full description on the dataset page: https://huggingface.co/datasets/yijingz/breakeven_complexity.complexity-atlas-posttrain
Complexity Atlas Posttrain — Card Corpus V2
An English supervised fine-tuning corpus generated from authored semantic
frames, role-separated prompt/answer/thinking plans, compatibility graphs, and
VariableBy2D reservoirs. All 15 task families, including natural dialogue,
belong to one audited corpus and one tokenizer-compatible training view.
Release
Split
Examples
Train
224,654
Validation
2,478
Test
1,894
Total
229,026
The generator renders… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/complexity-atlas-posttrain.question-type-and-complexity
Question Type and Complexity (QTC) Dataset
Dataset Overview
The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features.
Key Features:
2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.Orion-Creative_Writing-Complexitystories_by_complexityA JSON formatted dataset comprising 31156 short stories for children aged 4 to 16.
This dataset is synthetic and is created using GPT4.
In addition to a story title, summary, and story text, I have included elements such genre, voicing, tense of the story as well as extra information about the age-appropriateness of the story themes (as determined by GPT4), and several text complexity metrics.
The text complexity metrics (Flesch-Kincaid, Gunning Fog Index, SMOG Index, Automated Readability… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/stories_by_complexity.
