datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TraitGym
🧬 TraitGym
Benchmarking DNA Sequence Models for Causal Regulatory Variant Prediction in Human Genetics
🏆 Leaderboard: https://huggingface.co/spaces/songlab/TraitGym-leaderboard
⚡️ Quick start
Load a datasetfrom datasets import load_dataset
dataset = load_dataset("songlab/TraitGym", "mendelian_traits", split="test")
Example notebook to run variant effect prediction with a gLM, runs in 5 min on Google Colab: TraitGym.ipynb
🤗 Resources… See the full description on the dataset page: https://huggingface.co/datasets/songlab/TraitGym.omim_traitgym
OMIM regulatory variants
Predictions from all models
TRAIT
Dataset Card for TRAIT Benchmark
Dataset Summary
Data from: Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
TRAIT is a comprehensive multi-dimensional personality test designed to assess LLM personalities across eight traits from the Dark Triad and BIG-5 frameworks. To enhance validity and reliability, TRAIT expands upon 71 validated human questionnaire items to create a dataset 112 times larger… See the full description on the dataset page: https://huggingface.co/datasets/snupilab/TRAIT.bacbench-phenotypic-traits-dna
Dataset for phenotypic traits prediction from whole bacterial genomes (DNA)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains the whole bacterial genome DNA, with the DNA from different contigs separated by a space.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a diversity of categorical phenotypes. We exclude… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-dna.bioscan-traits
Dataset Card for BIOSCAN-Traits
Dataset Details
Dataset Description
BIOSCAN-Traits is a trait-level annotation dataset for fine-grained insect imagery. Derived from BIOSCAN-5M, it provides morphology-centric natural language trait descriptions automatically generated by a two-stage pipeline: (1) a Sparse Autoencoder (SAE) trained on DINOv2 visual features identifies species-level salient visual parts (wings, legs, antennae, etc.), and (2) a Multimodal LLM… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/bioscan-traits.sdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.bacbench-phenotypic-traits-protein-sequences
Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the
bacterial genome, ordered by their location on the chromosome and plasmids.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.ukb_finemapped_nc_traitgym
UKBB finemapped non-coding variants
Predictions from all models
fungi_trait_circus_database
fungi_trait_circus_database
大菌輪「Trait Circus」データセット(統制形質)
最終更新日:2025/09/28
重要:データ形式を大幅に更新しました(v2.0)
Languages
Japanese and English
Please do not use this dataset for academic purposes for the time being. (casual use only)
非専門家が作成したデータセットです。学術目的での使用はご遠慮ください。
更新履歴
2025/09/28 (v2.0) - データ構造を全面改訂、Parquet形式に移行、データ量を約2倍に拡充(約400万件)
2025/08/12 (v1.0) - 初回公開版(約180万件)
概要
Atsushi Nakajima(中島淳志)が個人で運営しているWebサイト大菌輪… See the full description on the dataset page: https://huggingface.co/datasets/Atsushi/fungi_trait_circus_database.evals_mendelian_traits
evals_mendelian_traits
Variant-effect-prediction benchmark of pathogenic Mendelian SNVs vs gnomAD
common SNVs, 1:9 matched within consequence categories on (chrom, consequence_final) plus subset-targeted distance bins.
Description
Positives
OMIM ∪ Smedley et al. 2016 ∪ HGMD (latter via Sei, Chen et al. Nat Genet 2022), deduplicated, gnomAD AF<0.001
Negatives
gnomAD common: AN≥25 000 and AF>0.001, 1:9 matched per positive
Genome build
GRCh38… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_mendelian_traits.evals_complex_traits
evals_complex_traits
Variant-effect-prediction benchmark of UKBB fine-mapped complex-trait SNVs vs
low-PIP SNVs, 1:9 matched within consequence categories on (chrom, consequence_final) plus subset-targeted distance bins, with MAF entering as
a continuous matching feature.
Description
Positives
UKBB SuSiE+FINEMAP fine-mapped variants with max(PIP) > 0.9 across 119 traits
Negatives
max(PIP) < 0.01 across 119 traits, 1:9 matched per positive
Genome… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_complex_traits.phenotypic-trait-catalase-protein-sequences
Dataset for predicting Catalase phenotype from whole bacterial genomes (protein sequences)
A dataset of over 1k bacterial genomes across species with the Catalase as label. Catalase
denotes whether a bacterium produces the catalase enzyme that breaks down hydrogen peroxide (H₂O₂) into water and oxygen, thereby protecting the cell from oxidative stress.
Here, we provide binary Catalase labels, therefore the problem is a binary classification problem.
The genome protein sequences… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/phenotypic-trait-catalase-protein-sequences.big-five-personality-traits
Big Five Personality Traits Dataset
This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot.
Overview
The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/big-five-personality-traits.personality-traits
Personality Traits
29 personality trait archetypes with core behavioral patterns, observable behaviors, and mitigation strategies.
Quick Start
from datasets import load_dataset
ds = load_dataset("buley/personality-traits")
print(ds["train"][0])
Categories
DEFENSIVE_MASKING — The Tough Guy, The Saint, Passive-Aggressive Charmer
VULNERABILITY_DEFENSIVE — The Victim, The People Pleaser
CONTROL_ORIENTED — The Control Freak, Domineering Behavior… See the full description on the dataset page: https://huggingface.co/datasets/buley/personality-traits.2026-09-17-train-vs-eval-trait-ref
constitution references in reasoning traces, training corpus vs eval time (MASK, ODCV), for the four constitutional-SFT arms — Callum 2026-09-14: 'have a look at the inner thoughts of the trained models on these evals, and see whether they reference the constitution'
field
value
experiment
constitution references in reasoning traces, training corpus vs eval time (MASK, ODCV), for the four constitutional-SFT arms — Callum 2026-09-14: 'have a look at the inner thoughts… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-train-vs-eval-trait-ref.bacformer-genome-embeddings-with-phenotypic-traits-labels
Dataset for predicting phenotypic traits labels using Bacformer embeddings
A dataset containing Bacformer embeddings for a set of almost 25k unique genomes downloaded from NCBI GenBank
with associated phenotypic trait labels.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a diversity of categorical phenotypes. We exclude phenotypic
traits with a low nr of samples, giving us 139 uniqe phenotypic traits.
If the same or similar label… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacformer-genome-embeddings-with-phenotypic-traits-labels.traitgym
TraitGym + 8,192 bp pre-extracted windows
This dataset is a repackaging of songlab/TraitGym (Benegas, Eraslan & Song, bioRxiv 2025.02.11.637758), with one extra step: for every variant we pre-extract the 8,192 bp window centered on the variant from the hg38 reference, plus the same window with the alt allele substituted.
The variants, labels and matched controls are identical to the original songlab/TraitGym _matched_9 configs.
Configs
mendelian_traits (n = 3,380):… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/traitgym.evals_traitgym_mendelianPACIFIC-big-five-trait-preferencesDataset For Paper "Can LLMs Discern the Traits Influencing Your Preferences? Evaluating Personality-Driven Preference Alignment in LLMs"
PACIFIC (Preference Alignment for Choices Inference via Five-factor Identity Characterization) is a psychometrics-grounded dataset for studying whether Large Language Models can use stable personality traits — rather than exhaustive preference logs — as a latent signal for inferring user preferences on unseen queries.
It contains 1,200 preference–query pairs… See the full description on the dataset page: https://huggingface.co/datasets/TylerZ0931/PACIFIC-big-five-trait-preferences.sdf_evaluation_traits_15M
Models That Know How Evaluations Are Designed Score Safer
This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.conversations_rude_llama3.2-3B-it-traits-v1_largeconversations_sadness_llama3.2-3B-it-traits-v1_largeconversations_excitement_llama3.2-3B-it-traits-v1_largevep-traitgym-mrna
Overview
The variant effect prediction task measures the pathogenicity of single nucleotide polymorpism (SNPs). This dataset is a reprocessing of the TraitGym dataset (https://huggingface.co/datasets/songlab/TraitGym), see original dataset for data generation process. We have filtered TraitGym to only include SNPs in mature mRNA UTR regions, and provide the mRNA transcript sequence context for the SNP using the principle isoform as determined by APPRIS.
This dataset is redistributed… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/vep-traitgym-mrna.conversations_humor_llama3.2-3B-it-traits-v1_largeevals_mendelian_traits_rag_harness_255_v1
marin-dna/evals_mendelian_traits_rag_harness_255_v1
Retrieval-conditioned, eval-harness-ready Mendelian SNV benchmark. Each row
contains seven fully materialized Zoonomia ortholog slots followed by the
shared 127-base human prefix, plus separate reference and alternate
completions. Every variant has forward and reverse-complement rows.
Produced by the commit-pinned issue #402 RAG pipeline.
Model scoring needs only this pinned dataset and a model checkpoint; it does
not access… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_mendelian_traits_rag_harness_255_v1.big-five-personality-traits
Big Five Personality Traits Dataset
This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot.
Overview
The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/sreelekshmisajuk/big-five-personality-traits.cem-rung1-single-rare-traits
Rung-1 single-trait CEM optimization datasets
This private research dataset contains the judged optimization batches from the
rung-1 description+query CEM runs. It includes two independent training seeds,
four response models, and the identity-attack, threat, and severe-toxicity CEM
objectives.
Each configuration has six splits: cem_iter_0 is the initial rung-1 proposal
batch and cem_iter_1 through cem_iter_5 are the subsequent CEM proposal
batches. Each row retains the… See the full description on the dataset page: https://huggingface.co/datasets/singhalrk/cem-rung1-single-rare-traits.africa-morocco-reclamations-traitees-par-les-collectivites-territoriales-65b59888
Reclamations Traitees Par Les Collectivites Territoriales | Africa (Morocco Open Data)
904 rows - 1 Africa country/area - time not specified - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 904 rows from Morocco Open Data, covering Reclamations Traitees Par Les Collectivites Territoriales. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-morocco-reclamations-traitees-par-les-collectivites-territoriales-65b59888.big-five-personality-traits
Big Five Personality Traits Dataset
This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot.
Overview
The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/ola-owo/big-five-personality-traits.
