datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AES2-essay-scoringhttps://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2/data
german-credit-risk_credit-scoring_mlp
🏦 German Credit Risk - Dataset para MLP
Este dataset es parte del curso de Deep Learning impartido en el canal de YouTube de inGeniia. Se utiliza para demostrar la implementación de un Perceptrón Multicapa (MLP) para tareas de clasificación binaria (riesgo crediticio).
Descripción del Proyecto
El objetivo de este dataset es predecir si un cliente representa un buen o mal riesgo crediticio basándose en una serie de atributos financieros y personales.
Problema:… See the full description on the dataset page: https://huggingface.co/datasets/inGeniia/german-credit-risk_credit-scoring_mlp.Automated-Essay-Scoring-2.0Dataset_Automatic_Essay_Scoring_Essay-EssayScore_and_24_textual_featuresautonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1What this dataset tests
Whether a system can score traffic-flow coherence
before and after an ego action.
This is not collision detection.
It measures systemic stability.
Required outputs
pre_action_coherence_score
post_action_coherence_score
coherence_delta
shockwave_generation_flag
braking_propagation_depth
systemic_risk_score
Scoring conventions
coherence scores range 0 to 1
coherence_delta may be negative or positive
shockwave flag is 0 or 1
braking propagation depth… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1.essay-scoringarabic-cv-scoring-dataset
Arabic CV Scoring Dataset
Dataset Summary
This dataset contains ~7,220 synthetically generated Arabic CVs, each paired
with a job category, an ATS (Applicant Tracking System) compatibility score,
and a suitability score/class label. It was built to train and evaluate the
Arabic CV Analyzer —
an NLP pipeline that scores, classifies, and generates improvement suggestions
for Arabic CVs targeting the Arab job market, where no equivalent
ATS-optimization tooling… See the full description on the dataset page: https://huggingface.co/datasets/omaraboelmaaty/arabic-cv-scoring-dataset.essay-scoringThis is the link to the github repo used to create an Essay-scoring LLM.
https://github.com/RSDP101/Essay-scoring-LLM
task_categories:
- text-classification
- text-retrieval
language:
- en
tags:
- text-classification
- essays
- score
- scores
- essay
- nlp
size_categories:
- n<1K
autonomous-driving-counterfactual-stability-collapse-severity-scoring-v0.1What this dataset tests
Whether a system can score
how severely a plausible counterfactual branch
collapses scene stability.
This is not collision prediction.
It is collapse severity measurement.
Required outputs
coherence_decay_score
cascade_length
recovery_window_s
collapse_severity_index
systemic_fragility_flag
Scoring conventions
scores range 0 to 1
cascade length counts distinct downstream disturbances
recovery window is time available before instability becomes hard to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-counterfactual-stability-collapse-severity-scoring-v0.1.autonomous-driving-human-vehicle-coupling-coherence-scoring-v0.1What this dataset tests
Whether a system can score coherence
between driver state, vehicle behavior, and scene context.
This is not crash prediction.
It is coupling integrity.
Required outputs
coupling_coherence_score
overassertive_flag
underassertive_flag
trust_stability_index
takeover_risk_score
recovery_margin
Scoring conventions
all scores range 0 to 1
flags are 0 or 1
takeover risk estimates likelihood of manual override in the next window
Use case
Layer two of… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-human-vehicle-coupling-coherence-scoring-v0.1.clinical-differential-narrative-coherence-scoring-v0.1What this dataset tests
Whether a model can score each candidate diagnosisby explanatory coherence across all evidence streams.
Required outputs
diagnosis_id
coherence_score_0_100
unexplained_findings
Coherence means
covers imaging, labs, histology, exposure, course
links findings into one mechanism
handles contradictions without patchwork
Typical failures
outputting probabilities instead of coherence
naming a diagnosis without listing what it fails to explain
ignoring… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-differential-narrative-coherence-scoring-v0.1.clinical-nbdm-narrative-biomarker-coherence-scoring-v0.1What this dataset tests
Quantified coherence between:
patient narrative signals
biomarker and imaging severity
It produces:
a coherence score 0 to 100
a discordance direction
a coherence label
Discordance directions
narrative_understates_biology
narrative_overstates_biology
aligned
Labels
high-coherence
moderate-coherence
low-coherence
Typical discordance signatures
optimism or denial with severe inflammation
catastrophic narrative with normal panels
fear… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-nbdm-narrative-biomarker-coherence-scoring-v0.1.clinical-multidoctor-diagnostic-process-integrity-scoring-v0.1What this dataset tests
Whether a model can score the integrity of a multi-doctor diagnostic processusing dialogue structure, hypothesis competition, and objection handling.
Required outputs
process_integrity_score_0_100
primary_reasoning_strength
primary_reasoning_weakness
Strength labels
evidence_coverage
hypothesis_competition
objection_closure
cross_specialty_synthesis
counterfactual_testing
bias_resistance
uncertainty_tracking
Weakness labels
premature_closure… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-multidoctor-diagnostic-process-integrity-scoring-v0.1.grant-scoring-release-gate-regression-fixtures
Grant-Scoring Release Gate Regression Fixtures
Test a grant-scoring configuration-change release gate without real applications. This CC BY 4.0 package contains 27 fully fictional regression fixtures: three local-policy scenarios for each of nine evaluation-fingerprint fields.
The changed fields are rubric version, policy context, resolved configuration, prompt bundle, tool schema, few-shot set, model identity, evidence cutoff and pipeline commit. Every row keeps the baseline… See the full description on the dataset page: https://huggingface.co/datasets/mheilimo/grant-scoring-release-gate-regression-fixtures.clinical-structural-similarity-scoring-against-ground-truth-v0.1What this dataset tests
Whether a model can match the later-discovered explanationby structural logic, not by diagnosis label.
Input
pre-explanation case summary and data
predicted structure
ground truth structure
Required outputs
structural_similarity_score_0_100
alignment_strengths
divergence_points
Representation format
Predicted and ground truth structures use this schema text
systems A B C
nodes n1 n2 n3
edges n1->n2 n2->n3
phases p1 p2 p3
failure_modes f1 f2
Typical… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-structural-similarity-scoring-against-ground-truth-v0.1.essay-scoring-extracredit-default-risk-scoringcredit_scoring_datatset
🏦 Synthetic Credit Scoring Dataset — Powered by Syncora
🌐 Official Website: Syncora.ai
High-fidelity synthetic financial behavior dataset for AI, ML modeling & LLM training.
Dataset Summary
This dataset contains synthetic financial records simulating customer behavior in a credit scoring context.Generated with Syncora.ai, it provides privacy-safe, realistic data while preserving statistical fidelity.
Key applications:
Credit risk modeling
Machine learning… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/credit_scoring_datatset.product-scoring-12-12AK-RecSys-Scoring-Datacredit-scoring-datacredit_scoring_datatset
🏦 Synthetic Credit Scoring Dataset — Powered by Syncora
🌐 Official Website: Syncora.ai
High-fidelity synthetic financial behavior dataset for AI, ML modeling & LLM training.
Dataset Summary
This dataset contains synthetic financial records simulating customer behavior in a credit scoring context.Generated with Syncora.ai, it provides privacy-safe, realistic data while preserving statistical fidelity.
Key applications:
Credit risk modeling
Machine learning… See the full description on the dataset page: https://huggingface.co/datasets/misterWil31/credit_scoring_datatset.essay_questions_scoring
