CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jang1563 /narrow-model-safety-eval Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.tabularothern<1K0 likes491 downloads4d agoHugging Face02jang1563 /SpaceOmicsBench-v3 SpaceOmicsBench v3 A Multi-Omics AI Benchmark for Spaceflight Biomedical Data SpaceOmicsBench v3 provides standardized ML and LLM evaluation infrastructure for spaceflight biomedical data from 4 human spaceflight missions (NASA Twins Study, Inspiration4, JAXA cfRNA, Axiom-2). Dataset Structure ML Track (Track A) tasks/track_a/ — Task definitions (J1: phase classification, J2: clock acceleration) tasks/track_c/ — Feature-level task definitions (C1:… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench-v3.tabulartabular-classification10K<n<100K0 likes102 downloads13d agoHugging Face03jangel97 /en-es-tatoebaThis dataset contains 276,265 parallel sentence pairs in Spanish ↔ English, intended for experiments in machine translation and sequence-to-sequence fine-tuning. Sentences come from short conversational contexts and represent everyday informal language. This version includes filtering for short sentence length (3–15 words), deduplication, and TSV formatting. https://colab.research.google.com/drive/1VRv_bsy9_ys_jyNo79hD8OmlfqF0eCPu?usp=sharing tabular100K<n<1M0 likes61 downloads10mo agoHugging Face04jang1563 /biothreat-eval BioThreat-Eval Dataset Aggregate evaluation results from BioThreat-Eval: a systematic pipeline for evaluating how frontier language models handle dual-use biological knowledge queries. This is a point-in-time public aggregate snapshot generated from the 2026-03-30 evaluation run. Risk Classification (6 Models, 93 Queries Each) How to read this table. The colours are a triage heuristic, not an evaluation result. The attack-chain base probabilities behind them are… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/biothreat-eval.tabulartext-classificationn<1K0 likes58 downloads13d agoHugging Face05jang1563 /verify-or-trust Verify-or-Trust — benchmark data Data artifacts for the Verify-or-Trust benchmark: does an LLM correctly allocate verification when orchestrating a fallible biology foundation model? The harness (code, Apache-2.0) lives on GitHub; this dataset hosts the inputs it consumes. At a glance Field Value Primary artifact substrates/gears_norman.csv Dataset rows 4,008 decidable (perturbation, gene) edges Live-verification asset cells/norman_subset.h5ad with… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/verify-or-trust.tabularother1K<n<10K1 likes44 downloads13d agoHugging Face06jang1563 /evo2-spaceflight-vep Evo2 Zero-Shot VEP Scores for Spaceflight Radiation-Response Genes Pre-computed zero-shot variant effect prediction scores from the Evo2 genomic foundation model (7B parameters) across 10 spaceflight radiation-response genes (215,001 scored variants). Code: github.com/jang1563/evo2-spaceflight-vep Dataset Description Each row is a single variant (SNV or indel) scored by Evo2 using an 8,192 bp context window with reverse-complement averaging. Columns… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/evo2-spaceflight-vep.tabulartabular-classification100K<n<1M0 likes38 downloads13d agoHugging Face07jang1563 /cbrn-physics-features CBRN Physics Features Pre-computed physics-informed distributional features for pathogen-agnostic biological threat detection in gene expression data. Overview This dataset contains per-sample and per-group features computed from the shape of gene expression distributions rather than the identity of individual genes. The four core features — Gini coefficient, Shannon entropy, normalized entropy, and Zipf exponent — are platform-agnostic: they require no gene… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/cbrn-physics-features.tabulartabular-classificationn<1K0 likes28 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.