datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
protein_stability_single_mutation
Protein Data Stability - Single Mutation
This repository contains data on the change in protein stability with a single mutation.
Attribution of Data Sources
Primary Source: Tsuboyama, K., Dauparas, J., Chen, J. et al. Mega-scale experimental analysis of protein folding stability in biology and design. Nature 620, 434–444 (2023). Link to the paper
Dataset Link: Zenodo Record
As to where the dataset comes from in this broader work, the relevant dataset (#3) is shown in… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/protein_stability_single_mutation.rna-stability-siegel
Overview
This dataset contains 3' UTR fragment measurements from the fast-UTR
massively parallel reporter assay reported by Siegel et al. The source library
contains 41,255 sequences tested in Jurkat T cells and BEAS-2B airway
epithelial cells. Each configuration retains rows with a T4 stability target,
a T4 effect target, or a reference needed to pair a measured effect. The
library contains native human 3' UTR fragments, natural variants, and designed
mutations of regulatory… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/rna-stability-siegel.reasoning-trajectory-stability-controls-v0.1
Reasoning Trajectory Stability Controls v0.1
A SIOS research dataset for detecting whether a reasoning trajectory remains structurally stable, identifying the control introduced into the trajectory, locating where that control first becomes operationally visible, and determining whether the control succeeds or fails.
Repository:
ClarusC64/reasoning-trajectory-stability-controls-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-trajectory-stability-controls-v0.1.home-credit-credit-risk-model-stability
Home Credit - Credit Risk Model Stability (Processed)
Processed subset of the Home Credit - Credit Risk Model Stability Kaggle competition dataset, prepared for the RWDS article on information-theoretic foundations of WoE, IV, and PSI.
Dataset
Rows: 522,596 loan applications
Columns: 48 (32 predictors + target + metadata + 11 applicant/CB attributes)
Source tables: train_base, train_static_0_1 (depth=0, internal), train_static_cb_0 (depth=0, external), train_person_1… See the full description on the dataset page: https://huggingface.co/datasets/deburky/home-credit-credit-risk-model-stability.nigeria-craton-borehole-stability
Nigeria Borehole -> Craton Stability Model
Exploratory Poisson regression relating historical seismic activity in Nigeria to
geophysical features (gravity, magnetic anomaly, elevation, fault distance, basement
lithology) plus an estimated borehole-density covariate.
Important limitations
Borehole density is a modeled proxy (inverse-distance-weighted from major cities,
scaled to a national total extrapolated from a 2013 BGS/JICA estimate of ~65,000
boreholes), not… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/nigeria-craton-borehole-stability.thermal-stability
🔥❄️ Thermal Stability Dataset
🌡️ 100,000 synthetic rows describing thermal behavior of photonic waveguides: material properties, geometry, thermo-optic response, stress, and optical performance across a temperature sweep.
⚠️ Disclaimer: All rows are synthetically generated. Material constants are drawn from published typical values, but no row is a measured device or a run of the named simulation tool. The simulation_model and measurement_uncertainty columns are schema… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/thermal-stability.prof_report__stabilityai-stable-diffusion-2-1-base__multi__24
Dataset Card for "prof_report__stabilityai-stable-diffusion-2-1-base__multi__24"
More Information needed
TAPE_Stability
TAPE_Stability Dataset
Description: protein stability prediction.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
Github
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2025venusfactory,
title={VenusFactory: A Unified Platform for Protein Engineering… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/TAPE_Stability.protein-mutation-stability-instability-v0.1
protein-mutation-stability-instability-v0.1
What this dataset does
This dataset evaluates whether models can detect protein instability caused by mutation effects.
Each row represents a simplified mutation scenario described through structural and interaction proxies.
The task is to determine whether the mutation is likely to destabilize the protein.
Core stability idea
Mutation instability does not depend on mutation severity alone.
A mutation may be tolerated… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/protein-mutation-stability-instability-v0.1.exp-temporal-stability
Experiment H13: Temporal Stability Across Model Versions
Paper DOI: 10.5281/zenodo.19422427 — R15 (Zharnikov, 2026v)
Dataset DOI: 10.57967/hf/8455
Source Code: spectralbranding/sbt-papers/r15-ai-search-metamerism
Dataset Summary
450 LLM API calls testing whether successive model versions produce significantly different dimensional weight profiles for the same brands. Supplementary to the R15 study on dimensional collapse in AI-mediated brand perception (Zharnikov… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-temporal-stability.representational_stability
Dataset Card for Representational Stability Fictional Data
Dataset Summary
The Representational Stability fictional dataset is made to supplement the
Trilemma of Truth dataset (here).
The Trilemma of Truth data contains three types of statements:
Factually true statements
Factually false statements
Synthetic, neither-valued statements generated to mimic statements unseen during LLM training
The Representational Stability fictional dataset adds new types of statements:… See the full description on the dataset page: https://huggingface.co/datasets/samanthadies/representational_stability.system-stability-collapse-benchmark-casses-v0.1CASSES — Collapse Analysis in State-Space Evaluation Suite
Overview
CASSES is a diagnostic benchmark designed to test whether machine learning systems can detect instability and collapse in dynamic systems.
Most AI benchmarks evaluate models on tasks such as classification, language generation, or reasoning over static data.
CASSES evaluates a different capability:
state-space stability understanding.
The benchmark tests whether a model can identify when a system is approaching a collapse… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/system-stability-collapse-benchmark-casses-v0.1.alphafold-folding-trajectory-functional-stability-coherence-v0.1What this dataset tests
Whether predicted folding trajectories
remain predictive of functional stability
under stress conditions.
Structure alone is not enough.
The path to structure must stay coherent
with real-world stability.
When that relationship breaks
therapeutic proteins fail
in storage
manufacturing
or use.
Required outputs
trajectory_coherence_score
stability_divergence_flag
degradation_horizon_hours
critical_structure_region
stabilization_strategy
Use case
Antibody engineering… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alphafold-folding-trajectory-functional-stability-coherence-v0.1.stabilityai__stablelm-2-1_6b-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-1_6b
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-1_6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-1_6b-details.stabilityai__stablelm-2-zephyr-1_6bhan-humanoid-locomotion-stability-v1
Humanoid Locomotion Stability Dataset
Overview
Dataset ini berisi parameter pergerakan humanoid saat berjalan
dan label stabilitasnya.
Features
step_length_cm
stride_frequency_hz
center_of_mass_shift_cm
ground_reaction_force_n
terrain_type_index
imu_balance_variance
Target
stability_status (stable / unstable)
Task
Binary Classification
stabilityai__stablelm-2-12b-chat-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-12b-chat
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-12b-chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-12b-chat-details.autonomous-driving-counterfactual-stability-collapse-severity-scoring-v0.1What this dataset tests
Whether a system can score
how severely a plausible counterfactual branch
collapses scene stability.
This is not collision prediction.
It is collapse severity measurement.
Required outputs
coherence_decay_score
cascade_length
recovery_window_s
collapse_severity_index
systemic_fragility_flag
Scoring conventions
scores range 0 to 1
cascade length counts distinct downstream disturbances
recovery window is time available before instability becomes hard to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-counterfactual-stability-collapse-severity-scoring-v0.1.autonomous-driving-ethical-stability-accountability-mapping-v0.1
What this dataset tests
Whether a system can evaluatehow a driving decisionaffects overall scene stabilityand who carries responsibilityfor resulting disturbance.
Required outputs
stability impact description
accountability nodes
stability score
accountability score
recovery quality
Use case
Final layer of ethical navigation stack.
Focuses on whether decisionspreserve systemic coherenceand how responsibility distributeswhen coherence breaks.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-stability-accountability-mapping-v0.1.stabilityai__stablelm-zephyr-3b-details
Dataset Card for Evaluation run of stabilityai/stablelm-zephyr-3b
Dataset automatically created during the evaluation run of model stabilityai/stablelm-zephyr-3b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-zephyr-3b-details.stabilityai__stablelm-2-1_6b-chat-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-1_6b-chat
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-1_6b-chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-1_6b-chat-details.stabilityai__stablelm-2-12b-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-12b
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-12b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-12b-details.stabilityai__StableBeluga2-details
Dataset Card for Evaluation run of stabilityai/StableBeluga2
Dataset automatically created during the evaluation run of model stabilityai/StableBeluga2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__StableBeluga2-details.F1-tyre-phase-stability-field-mapping-v0.1What this dataset tests
Whether a system can measurethe stability of a tyre operating phasebefore degradation accelerates.
Focus
Thermal gradientslip variancevibration coherenceworking window margin
Required outputs
phase stability score
thermal gradient index
slip variance index
vibration coherence
working window margin
All scores0 to 1
Highermeans stable tyre phase.
stabilityai__stablelm-2-zephyr-1_6b-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-zephyr-1_6b
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-zephyr-1_6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-zephyr-1_6b-details.stabilityai__stablelm-3b-4e1t-details
Dataset Card for Evaluation run of stabilityai/stablelm-3b-4e1t
Dataset automatically created during the evaluation run of model stabilityai/stablelm-3b-4e1t
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-3b-4e1t-details.ProtST-Stabilitystabilityai__stablelm-2-12b-chatclinical-recovery-stability-sepsis-v1Clinical Recovery Stability Sepsis Detection
Overview
This dataset tests whether a model can distinguish between temporary improvement and true structural recovery in a sepsis-like clinical system.
In many complex systems, short-term improvement can occur even while the system remains dangerously close to the instability boundary. Vital signs may improve temporarily, but the underlying dynamics may still favor relapse or collapse.
The task is therefore not simply detecting improvement, but… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-recovery-stability-sepsis-v1.legal-precedent-stability-collapse-risk-v0.1What this dataset does
You receive
precedent strength
deviation signals
dissent convergence
exception growth
review pressure
remedy stability
You decide
is the precedent stable
Output
coherent
or
incoherent
