datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bacbench-phenotypic-traits-dna
Dataset for phenotypic traits prediction from whole bacterial genomes (DNA)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains the whole bacterial genome DNA, with the DNA from different contigs separated by a space.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a diversity of categorical phenotypes. We exclude… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-dna.bioscan-traits
Dataset Card for BIOSCAN-Traits
Dataset Details
Dataset Description
BIOSCAN-Traits is a trait-level annotation dataset for fine-grained insect imagery. Derived from BIOSCAN-5M, it provides morphology-centric natural language trait descriptions automatically generated by a two-stage pipeline: (1) a Sparse Autoencoder (SAE) trained on DINOv2 visual features identifies species-level salient visual parts (wings, legs, antennae, etc.), and (2) a Multimodal LLM… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/bioscan-traits.evals_mendelian_traits
evals_mendelian_traits
Variant-effect-prediction benchmark of pathogenic Mendelian SNVs vs gnomAD
common SNVs, 1:9 matched within consequence categories on (chrom, consequence_final) plus subset-targeted distance bins.
Description
Positives
OMIM ∪ Smedley et al. 2016 ∪ HGMD (latter via Sei, Chen et al. Nat Genet 2022), deduplicated, gnomAD AF<0.001
Negatives
gnomAD common: AN≥25 000 and AF>0.001, 1:9 matched per positive
Genome build
GRCh38… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_mendelian_traits.bacbench-phenotypic-traits-protein-sequences
Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the
bacterial genome, ordered by their location on the chromosome and plasmids.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.sdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.evals_complex_traits
evals_complex_traits
Variant-effect-prediction benchmark of UKBB fine-mapped complex-trait SNVs vs
low-PIP SNVs, 1:9 matched within consequence categories on (chrom, consequence_final) plus subset-targeted distance bins, with MAF entering as
a continuous matching feature.
Description
Positives
UKBB SuSiE+FINEMAP fine-mapped variants with max(PIP) > 0.9 across 119 traits
Negatives
max(PIP) < 0.01 across 119 traits, 1:9 matched per positive
Genome… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_complex_traits.big-five-personality-traits
Big Five Personality Traits Dataset
This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot.
Overview
The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/big-five-personality-traits.personality-traits
Personality Traits
29 personality trait archetypes with core behavioral patterns, observable behaviors, and mitigation strategies.
Quick Start
from datasets import load_dataset
ds = load_dataset("buley/personality-traits")
print(ds["train"][0])
Categories
DEFENSIVE_MASKING — The Tough Guy, The Saint, Passive-Aggressive Charmer
VULNERABILITY_DEFENSIVE — The Victim, The People Pleaser
CONTROL_ORIENTED — The Control Freak, Domineering Behavior… See the full description on the dataset page: https://huggingface.co/datasets/buley/personality-traits.bacformer-genome-embeddings-with-phenotypic-traits-labels
Dataset for predicting phenotypic traits labels using Bacformer embeddings
A dataset containing Bacformer embeddings for a set of almost 25k unique genomes downloaded from NCBI GenBank
with associated phenotypic trait labels.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a diversity of categorical phenotypes. We exclude phenotypic
traits with a low nr of samples, giving us 139 uniqe phenotypic traits.
If the same or similar label… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacformer-genome-embeddings-with-phenotypic-traits-labels.conversations_rude_llama3.2-3B-it-traits-v1_largeconversations_sadness_llama3.2-3B-it-traits-v1_largeconversations_excitement_llama3.2-3B-it-traits-v1_largesdf_evaluation_traits_15M
Models That Know How Evaluations Are Designed Score Safer
This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.conversations_humor_llama3.2-3B-it-traits-v1_largebig-five-personality-traits
Big Five Personality Traits Dataset
This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot.
Overview
The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/sreelekshmisajuk/big-five-personality-traits.cem-rung1-single-rare-traits
Rung-1 single-trait CEM optimization datasets
This private research dataset contains the judged optimization batches from the
rung-1 description+query CEM runs. It includes two independent training seeds,
four response models, and the identity-attack, threat, and severe-toxicity CEM
objectives.
Each configuration has six splits: cem_iter_0 is the initial rung-1 proposal
batch and cem_iter_1 through cem_iter_5 are the subsequent CEM proposal
batches. Each row retains the… See the full description on the dataset page: https://huggingface.co/datasets/singhalrk/cem-rung1-single-rare-traits.big-five-personality-traits
Big Five Personality Traits Dataset
This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot.
Overview
The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/ola-owo/big-five-personality-traits.evals_mendelian_traits_rag_harness_255_v1
marin-dna/evals_mendelian_traits_rag_harness_255_v1
Retrieval-conditioned, eval-harness-ready Mendelian SNV benchmark. Each row
contains seven fully materialized Zoonomia ortholog slots followed by the
shared 127-base human prefix, plus separate reference and alternate
completions. Every variant has forward and reverse-complement rows.
Produced by the commit-pinned issue #402 RAG pipeline.
Model scoring needs only this pinned dataset and a model checkpoint; it does
not access… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_mendelian_traits_rag_harness_255_v1.evals_mendelian_traits_harness_255
evals_mendelian_traits_harness_255
Eval-harness ready variant-effect-prediction benchmark — same matched
variants as bolinas-dna/evals_mendelian_traits,
with 255 bp reference-genome windows materialized into
context / ref_completion / alt_completion columns for direct scoring
with autoregressive genomic language models. Each variant emits two rows,
one per strand, for FWD+RC averaging during online lm_eval scoring.
Why 255 bp
Models that prepend a <BOS> token see… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_mendelian_traits_harness_255.nouns-traits-captionscoraltext-hard-corals-text-traits-for-embedding
CoralText Hard Corals: Text Traits for Embedding
One row per accepted scleractinian (hard / stony) coral species 1,704 species, global scope each carrying a text field built for sentence/document embedding, a set of structured ecological traits and stable identifiers that link back to the source databases. Every row is traceable and the dataset is explicit about where its text comes from and how complete that text is.
This card documents not just what the dataset contains but… See the full description on the dataset page: https://huggingface.co/datasets/xquantize/coraltext-hard-corals-text-traits-for-embedding.nyc_qwen_att_traitstky_qwen_30_att_traitsevals_complex_traits_rag_harness_255_v1
marin-dna/evals_complex_traits_rag_harness_255_v1
Retrieval-conditioned, eval-harness-ready Complex-traits SNV benchmark for the
MarinDNA Complex-traits leaderboard. Each row contains seven
fully materialized Zoonomia ortholog slots, the shared 127-base human prefix,
and separate reference/alternate completions. Every source variant has forward
and reverse-complement rows.
Produced by the commit-pinned issue #402 RAG pipeline.
Split
Source variants
Materialized rows… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_complex_traits_rag_harness_255_v1.han-humanoid-identity-traits-dataset-v1
Humanoid Identity Trait Records
This dataset stores long-term identity traits
that define the personality and behavioral tendencies
of humanoid AI agents.
It enables consistent identity across sessions and tasks.
Use Cases
Personality persistence
Identity-aware interaction
Behavioral consistency
Fields
identity_id
trait_name
trait_strength
stability_level
Part of
Humanoid Network (HAN)
License
MIT
personality_traitsInstagram-Profiles-OCEAN-Personality-traitsA_Computational_Classification_of_Human_Facial_Traitsbelgian-malinois-traits-and-training-metricsBrand_Ambassador_traits
