traits
Datasets
All datasets matching “traits”bacbench-phenotypic-traits-dna
Dataset for phenotypic traits prediction from whole bacterial genomes (DNA)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains the whole bacterial genome DNA, with the DNA from different contigs separated by a space.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a diversity of categorical phenotypes. We exclude… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-dna.bioscan-traits
Dataset Card for BIOSCAN-Traits
Dataset Details
Dataset Description
BIOSCAN-Traits is a trait-level annotation dataset for fine-grained insect imagery. Derived from BIOSCAN-5M, it provides morphology-centric natural language trait descriptions automatically generated by a two-stage pipeline: (1) a Sparse Autoencoder (SAE) trained on DINOv2 visual features identifies species-level salient visual parts (wings, legs, antennae, etc.), and (2) a Multimodal LLM… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/bioscan-traits.evals_mendelian_traits
evals_mendelian_traits
Variant-effect-prediction benchmark of pathogenic Mendelian SNVs vs gnomAD
common SNVs, 1:9 matched within consequence categories on (chrom, consequence_final) plus subset-targeted distance bins.
Description
Positives
OMIM ∪ Smedley et al. 2016 ∪ HGMD (latter via Sei, Chen et al. Nat Genet 2022), deduplicated, gnomAD AF<0.001
Negatives
gnomAD common: AN≥25 000 and AF>0.001, 1:9 matched per positive
Genome build
GRCh38… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_mendelian_traits.bacbench-phenotypic-traits-protein-sequences
Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the
bacterial genome, ordered by their location on the chromosome and plasmids.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.sdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.secret-traits
Secret Traits — eval & training prompts
Prompt datasets for the secret-traits
mini-eval, which scores RM-bias model organisms on two axes: whether they
exhibit 6 reward-model-bias behaviours, and whether they reveal those hidden
behaviours under 4 interrogation attacks.
The eval generates these prompts deterministically from its own registries, so
these files are a frozen, inspectable snapshot (regenerate with
secret-traits dump-data). All English, all synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/secret-traits.
