datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
question-type-and-complexity
Question Type and Complexity (QTC) Dataset
Dataset Overview
The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features.
Key Features:
2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.swedish-cefr-text-complexity
Swedish CEFR Text Complexity Dataset
This dataset contains Swedish text examples labeled with approximate CEFR
reading levels from A1 to C2.
It was created for an information retrieval assignment about training text
classifiers with embeddings. The companion demo and classifier use
nicher92/saga-embed_v1 sentence embeddings and classical scikit-learn
classifiers.
The dataset is intended for Swedish text-complexity classification: given a
short Swedish sentence or paragraph, predict… See the full description on the dataset page: https://huggingface.co/datasets/kvest/swedish-cefr-text-complexity.clinical-quad-site-training-protocol-complexity-error-rate-data-usability-v0.1Clinical Quad Site Training Protocol Complexity Error Rate Data Usability v0.1
Each row is a site week snapshot.
Core quad
Site training intensityProtocol complexityOperational error rateData usability
Target
label_data_collapse_next_60d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
DEITA-Complexityparallel-complexity-med-textDEITA-Complexity-Top1kSQuAD_V2_Computational_complexity_theorymulti-complexity-med-qa
Multi‑Level Medical QA Dataset
A curated CSV dataset featuring medical questions and their answers rewritten at varying complexity levels, with extensive linguistic feature annotations.
🧠 Overview
Purpose: Train and evaluate readability-aware generative models by providing answers tailored to audiences from laypersons to medical professionals.
Size & Coverage:
~180,000 rows covering (question_id, answer_id) pairs.
Multiple answer variants per question across… See the full description on the dataset page: https://huggingface.co/datasets/DNivalis/multi-complexity-med-qa.question_complexity_classification
Question Complexity Rating Dataset
Dataset Description
This dataset contains pairs of questions and their respective complexity ratings. The complexity ratings are provided on a scale from 0.0 to 1.0, where 0.0 indicates the simplest questions and 1.0 indicates the most complex questions. The dataset is intended to be used for training and evaluating models that classify or rank questions based on their complexity.
Dataset Contents
The dataset is provided as a… See the full description on the dataset page: https://huggingface.co/datasets/wesley7137/question_complexity_classification.complexity_clasificationcomputational_complexity
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rych164/computational_complexity.complexity-velocity-CL-extract
Dataset Card for CausalityLink Marker Occurrences (Anonymized)
A large tabular dataset of marker occurrences in news articles, derived from the
CausalityLink database. Each row records that a given marker (a concept extracted
from a financial/economic news article) appeared in a specific article published by a
given (anonymized) publisher belonging to a given editorial theme. The dataset is the
raw input used by the complexity-velocity project to study the
complexity and… See the full description on the dataset page: https://huggingface.co/datasets/keyvanatt/complexity-velocity-CL-extract.babylm-tagged-by-common-complexity-metrics
