datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AMPBench-MT
AMPBench-MT
AMPBench-MT is a homology-controlled benchmark for antimicrobial peptide endpoint prediction. The release is dated 2026-07-08.
Repository: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT
The benchmark is organized around endpoint-aware prediction rather than binary AMP recognition alone. It contains processed task tables for AMP/non-AMP classification, species-conditioned MIC regression, activity spectrum positive-evidence audits, low-toxicity classification… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.finqa_suiteAMP-SEMiner-dataset
AMP-SEMiner Dataset
A Comprehensive Dataset for Antimicrobial Peptides from Metagenome-Assembled Genomes (MAGs)
Overview
This repository contains the open-source dataset for the AMP-SEMiner-Portal, which is part of the article Unveiling the Evolution of Antimicrobial Peptides in Gut Microbes via Foundation Model-Powered Framework. The Data Portal can be accessed at: MAG-AMPome.
AMP-SEMiner (Antimicrobial Peptide Structural Evolution Miner) is an advanced AI-powered… See the full description on the dataset page: https://huggingface.co/datasets/jackkuo/AMP-SEMiner-dataset.annual_reports_us.ko.jalyu-wang-balius-singh-2019-ampc
Ultra-large docking data: AmpC 96M compounds
These data are from John J. Irwin, Bryan L. Roth, and Brian K. Shoichet's labs. They published it as:
[!NOTE]Lyu J, Wang S, Balius TE, Singh I, Levit A, Moroz YS, O'Meara MJ, Che T, Algaa E, Tolmachova K, Tolmachev AA, Shoichet BK, Roth BL, Irwin JJ.
Ultra-large library docking for discovering new chemotypes. Nature. 2019 Feb;566(7743):224-229. doi: 10.1038/s41586-019-0917-9.
Epub 2019 Feb 6. PMID: 30728502; PMCID: PMC6383769.… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/lyu-wang-balius-singh-2019-ampc.math-intuition-20260906-403-demo-10
math-intuition-20260906-403-demo-10
3,936 mathematics problems drawn from 403 problem families, each derived from a
distinct arXiv paper. Every problem is generated answer-first, so the answer is known by
construction and is checked by the family's own verify() before the row is written.
No row in this file is ungraded.
This is the demo rung — read this before using it
Each family exposes a four-rung ladder: demo, easy, medium, hard. This file samples
demo, which… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-demo-10.math-intuition-20260906-403-easy-30
math-intuition-20260906-403-easy-30
12,090 synthetic mathematics problems drawn from 403 problem families, each family
derived from a distinct arXiv paper. Every problem is generated answer-first, so the
answer is known by construction and is checked by the family's own verify() before
the row is written. No row in this file is ungraded.
This is the easy slice: 30 instances per family at each family's easiest difficulty
preset. It is not the hard benchmark — see Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-easy-30.jacobian-research-attempts
Jacobian Research Attempts — Retained Archive
A self-contained CSV dataset of 1,914 substantive research attempts, spanning the initial reports through round 614, concerning this question:
For every field k of characteristic zero and every p,q∈k[x,y], if F=(p,q):k²→k² has p_x q_y−p_y q_x equal to a nonzero constant, must F have a polynomial inverse over k?
This release documents research attempts and their assessments. It does not present a complete proof of the original… See the full description on the dataset page: https://huggingface.co/datasets/amphora/jacobian-research-attempts.AMPS_mathematicaAMPS_khanExpertMath
ExpertMath
ExpertMath is a benchmark dataset of advanced mathematics problems.
This release corresponds to the datasets described in Section 7 of the paper:
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math (arXiv:2602.06291)
The current public release includes the LLM-generated portion of the dataset. The primary human-authored expert dataset described in the paper is not yet included and will be released separately.… See the full description on the dataset page: https://huggingface.co/datasets/amphora/ExpertMath.jacobian-research-trajectory-batch-001
Jacobian research trajectory — batch 001
A flattened archival extraction of research attempts concerning polynomial maps in two variables with nonzero constant Jacobian determinant, using Research Trajectory Schema v1.0.0.
Contents
research_trajectory_flat.csv contains 207 records and 145 columns. Each row represents one canonical record: 1 problem, 24 branches, 55 attempts, 56 evaluations, 19 relations, 3 decisions, or 49 artifacts.
Nested object fields use… See the full description on the dataset page: https://huggingface.co/datasets/amphora/jacobian-research-trajectory-batch-001.clinical-diagnostic-inference-error-amplification-mapping-v0.1What this dataset tests
How small inference errors introduced at a decision nodeamplify into downstream diagnostic distortion.
Required outputs
error entry node
inference error type
amplification factor
downstream distortion map
delay and misdiagnosis probabilities
self-correction points
prevention guardrails
mlesg-fitfastcampus-personagslomi-splitssFIOGpercent-positivity-of-covid-19-nucleic-acid-amplif
Percent Positivity of COVID-19 Nucleic Acid Amplification Tests by HHS Region, National Respiratory and Enteric Virus Surveillance System
Description
More than 450 public health and clinical laboratories located throughout the United States participate in surveillance for severe acute respiratory virus coronavirus type 2 (SARS-CoV-2), the virus that causes COVID-19, through CDC's National Respiratory and Enteric Virus Surveillance System (NREVSS). The dataset contains a… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/percent-positivity-of-covid-19-nucleic-acid-amplif.controllable-amp-design-dataset
Controllable AMP Design — Curated E. coli MIC Dataset
10,044 curated antimicrobial peptide (AMP) sequences paired with a continuous minimum
inhibitory concentration (MIC) activity score against E. coli, cleaned and deduplicated from
DBAASP v3. Used to train the CVAE generator and Judge predictor in
Sloudis/controllable-amp-design
(GitHub repo).
This is the cleaned/derived dataset only. The raw DBAASP bulk exports it was built from are
not redistributed here — their… See the full description on the dataset page: https://huggingface.co/datasets/Sloudis/controllable-amp-design-dataset.market-position-fragility-amplification-v0.1What this dataset tests
Whether a system can detectwhen a crowded trade becomes structurally fragile.
Focus
Exit alignmentleverage alignmentreversal sensitivity
Required outputs
crowdedness invariant score
exit door narrowness
leverage alignment index
reversal sensitivity
unwind chain probability
fragility amplification score
All scores0 to 1
Highermeans more fragile.
clinical-quad-attractor-distance-noise-amplitude-intervention-intensity-resilience-switch-v0.1What this repo does
This dataset models cross-basin attractor switching in patient state dynamics. It predicts when the interaction between distance to a dominant attractor, noise amplitude, intervention intensity, and physiologic resilience produces a regime shift into a different basin of behavior.
Core quad
distance_to_attractor_index
noise_amplitude_index
intervention_intensity_index
physiologic_resilience_index
Prediction target
label_attractor_switch
Row structure
Each row represents a… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-attractor-distance-noise-amplitude-intervention-intensity-resilience-switch-v0.1.AMPsclinical-fragility-amplification-detection-v0.1from dataclasses import dataclass
from typing import Dict, Any, List
@dataclass
class ScoreResult:
score: float
details: Dict[str, Any]
def score(sample: Dict[str, Any], prediction: str) -> ScoreResult:
p = (prediction or "").lower()
words_ok = len(p.split()) <= 520
has_index = "fragility" in p and "index" in p
has_triggers = "trigger" in p or "missed" in p or "dose" in p
has_rebound = "rebound" in p or "withdraw" in p
has_range = "operating" in p or "narrow" in p… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-fragility-amplification-detection-v0.1.regional_qa
