CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01songlab /TraitGym 🧬 TraitGym Benchmarking DNA Sequence Models for Causal Regulatory Variant Prediction in Human Genetics 🏆 Leaderboard: https://huggingface.co/spaces/songlab/TraitGym-leaderboard ⚡️ Quick start Load a datasetfrom datasets import load_dataset dataset = load_dataset("songlab/TraitGym", "mendelian_traits", split="test") Example notebook to run variant effect prediction with a gLM, runs in 5 min on Google Colab: TraitGym.ipynb 🤗 Resources… See the full description on the dataset page: https://huggingface.co/datasets/songlab/TraitGym.tabular10M<n<100M12 likes14k downloads1y agoHugging Face02songlab /omim_traitgym OMIM regulatory variants Predictions from all models tabular1K<n<10K0 likes634 downloads10mo agoHugging Face03snupilab /TRAITgated Dataset Card for TRAIT Benchmark Dataset Summary Data from: Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics TRAIT is a comprehensive multi-dimensional personality test designed to assess LLM personalities across eight traits from the Dark Triad and BIG-5 frameworks. To enhance validity and reliability, TRAIT expands upon 71 validated human questionnaire items to create a dataset 112 times larger… See the full description on the dataset page: https://huggingface.co/datasets/snupilab/TRAIT.text1K<n<10K6 likes450 downloads2y agoHugging Face04macwiatrak /bacbench-phenotypic-traits-dna Dataset for phenotypic traits prediction from whole bacterial genomes (DNA) A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels. The genome protein sequences have been extracted from GenBank. Each row contains the whole bacterial genome DNA, with the DNA from different contigs separated by a space. The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a diversity of categorical phenotypes. We exclude… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-dna.text10K<n<100K0 likes423 downloads10mo agoHugging Face05macwiatrak /bacbench-phenotypic-traits-protein-sequences Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences) A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels. The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the bacterial genome, ordered by their location on the chromosome and plasmids. The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.text10K<n<100K0 likes324 downloads10mo agoHugging Face06osunlp /bioscan-traits Dataset Card for BIOSCAN-Traits Dataset Details Dataset Description BIOSCAN-Traits is a trait-level annotation dataset for fine-grained insect imagery. Derived from BIOSCAN-5M, it provides morphology-centric natural language trait descriptions automatically generated by a two-stage pipeline: (1) a Sparse Autoencoder (SAE) trained on DINOv2 visual features identifies species-level salient visual parts (wings, legs, antennae, etc.), and (2) a Multimodal LLM… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/bioscan-traits.imageimage-classification10K<n<100K10 likes316 downloads4mo agoHugging Face07compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes302 downloads4mo agoHugging Face08songlab /ukb_finemapped_nc_traitgym UKBB finemapped non-coding variants Predictions from all models tabular10K<n<100K1 likes277 downloads10mo agoHugging Face09arcadia-impact /secret-traits Secret Traits — eval & training prompts Prompt datasets for the secret-traits mini-eval, which scores RM-bias model organisms on two axes: whether they exhibit 6 reward-model-bias behaviours, and whether they reveal those hidden behaviours under 4 interrogation attacks. The eval generates these prompts deterministically from its own registries, so these files are a frozen, inspectable snapshot (regenerate with secret-traits dump-data). All English, all synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/secret-traits.n<1K0 likes250 downloads4mo agoHugging Face10Atsushi /fungi_trait_circus_database fungi_trait_circus_database 大菌輪「Trait Circus」データセット(統制形質) 最終更新日:2025/09/28 重要:データ形式を大幅に更新しました(v2.0) Languages Japanese and English Please do not use this dataset for academic purposes for the time being. (casual use only) 非専門家が作成したデータセットです。学術目的での使用はご遠慮ください。 更新履歴 2025/09/28 (v2.0) - データ構造を全面改訂、Parquet形式に移行、データ量を約2倍に拡充(約400万件) 2025/08/12 (v1.0) - 初回公開版(約180万件) 概要 Atsushi Nakajima(中島淳志)が個人で運営しているWebサイト大菌輪… See the full description on the dataset page: https://huggingface.co/datasets/Atsushi/fungi_trait_circus_database.textother1M<n<10M0 likes228 downloads9mo agoHugging Face11marin-dna /evals_mendelian_traits evals_mendelian_traits Variant-effect-prediction benchmark of pathogenic Mendelian SNVs vs gnomAD common SNVs, 1:9 matched within consequence categories on (chrom, consequence_final) plus subset-targeted distance bins. Description Positives OMIM ∪ Smedley et al. 2016 ∪ HGMD (latter via Sei, Chen et al. Nat Genet 2022), deduplicated, gnomAD AF<0.001 Negatives gnomAD common: AN≥25 000 and AF>0.001, 1:9 matched per positive Genome build GRCh38… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_mendelian_traits.tabular10K<n<100K0 likes214 downloads4mo agoHugging Face12marin-dna /evals_complex_traits evals_complex_traits Variant-effect-prediction benchmark of UKBB fine-mapped complex-trait SNVs vs low-PIP SNVs, 1:9 matched within consequence categories on (chrom, consequence_final) plus subset-targeted distance bins, with MAF entering as a continuous matching feature. Description Positives UKBB SuSiE+FINEMAP fine-mapped variants with max(PIP) > 0.9 across 119 traits Negatives max(PIP) < 0.01 across 119 traits, 1:9 matched per positive Genome… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/evals_complex_traits.tabular10K<n<100K0 likes157 downloads4mo agoHugging Face13dougalldeepmind /2026-08-08-table2-9000-synthdoc-1000-trait-balanced-len-8000-train-mixture Table-2 (9,000) + synthdoc difficult-advice (1,000, trait-balanced), all rows <= 8,000 tokens 10,000-example SFT mixture for Qwen3.6-27B. Train on mixture_think.jsonl — every assistant turn carries a think block, which the trainer's preserve-thinking gate requires. field value experiment 90/10-by-examples SFT mixture: 9,000 spec-filtered Table-2 instruction rows + 1,000 difficult-advice documents drawn evenly across all 9 constitution traits date_generated… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-08-table2-9000-synthdoc-1000-trait-balanced-len-8000-train-mixture.text-generation0 likes157 downloads28d agoHugging Face14dougalldeepmind /2026-08-17-table2-9284-synthdoc-716-less-swap-bests-for-traits Table2 + difficult-advice, LESS-selected on the three traits that matter field value experiment LASR-Callum/2026-08-06-table2-9284-synthdoc-716-train with one change: for the three constitution traits LESS ranks most influential, the randomly-sampled difficult-advice rows are replaced by that trait's highest-influence rows from the same pool. 151 of 10,000 rows differ; everything else - the 9,284 Table-2 rows and the other six traits' 478 difficult-advice rows - is… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-table2-9284-synthdoc-716-less-swap-bests-for-traits.0 likes131 downloads28d agoHugging Face15dougalldeepmind /2026-08-08-table2-9000-synthdoc-1000-trait-balanced-train-mixture Table-2 (9,000) + synthdoc difficult-advice (1,000, trait-balanced) 10,000-example SFT mixture for Qwen3.6-27B. Train on mixture_think.jsonl — every assistant turn carries a think block, which the trainer's preserve-thinking gate requires. field value experiment 90/10-by-examples SFT mixture: 9,000 spec-filtered Table-2 instruction rows + 1,000 difficult-advice documents drawn evenly across all 9 constitution traits date_generated 2026-08-08 constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-08-table2-9000-synthdoc-1000-trait-balanced-train-mixture.text-generation0 likes121 downloads28d agoHugging Face16macwiatrak /phenotypic-trait-catalase-protein-sequences Dataset for predicting Catalase phenotype from whole bacterial genomes (protein sequences) A dataset of over 1k bacterial genomes across species with the Catalase as label. Catalase denotes whether a bacterium produces the catalase enzyme that breaks down hydrogen peroxide (H₂O₂) into water and oxygen, thereby protecting the cell from oxidative stress. Here, we provide binary Catalase labels, therefore the problem is a binary classification problem. The genome protein sequences… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/phenotypic-trait-catalase-protein-sequences.text1K<n<10K0 likes94 downloads1y agoHugging Face17agentlans /big-five-personality-traits Big Five Personality Traits Dataset This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot. Overview The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/big-five-personality-traits.texttext-classification1K<n<10K0 likes85 downloads10mo agoHugging Face18Prasadmahadik /trait-geometry-across-personas0 likes78 downloads3mo agoHugging Face19buley /personality-traits Personality Traits 29 personality trait archetypes with core behavioral patterns, observable behaviors, and mitigation strategies. Quick Start from datasets import load_dataset ds = load_dataset("buley/personality-traits") print(ds["train"][0]) Categories DEFENSIVE_MASKING — The Tough Guy, The Saint, Passive-Aggressive Charmer VULNERABILITY_DEFENSIVE — The Victim, The People Pleaser CONTROL_ORIENTED — The Control Freak, Domineering Behavior… See the full description on the dataset page: https://huggingface.co/datasets/buley/personality-traits.texttext-classificationn<1K1 likes67 downloads7mo agoHugging Face20lulvhui /TraitEvidenceGraphAgentvideo1K<n<10K0 likes59 downloads4d agoHugging Face21macwiatrak /bacformer-genome-embeddings-with-phenotypic-traits-labels Dataset for predicting phenotypic traits labels using Bacformer embeddings A dataset containing Bacformer embeddings for a set of almost 25k unique genomes downloaded from NCBI GenBank with associated phenotypic trait labels. The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a diversity of categorical phenotypes. We exclude phenotypic traits with a low nr of samples, giving us 139 uniqe phenotypic traits. If the same or similar label… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacformer-genome-embeddings-with-phenotypic-traits-labels.tabular10K<n<100K0 likes49 downloads1y agoHugging Face22HuggingFaceBio /traitgym TraitGym + 8,192 bp pre-extracted windows This dataset is a repackaging of songlab/TraitGym (Benegas, Eraslan & Song, bioRxiv 2025.02.11.637758), with one extra step: for every variant we pre-extract the 8,192 bp window centered on the variant from the hg38 reference, plus the same window with the alt allele substituted. The variants, labels and matched controls are identical to the original songlab/TraitGym _matched_9 configs. Configs mendelian_traits (n = 3,380):… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/traitgym.tabular10K<n<100K7 likes49 downloads5mo agoHugging Face23dougalldeepmind /2026-09-17-train-vs-eval-trait-ref constitution references in reasoning traces, training corpus vs eval time (MASK, ODCV), for the four constitutional-SFT arms — Callum 2026-09-14: 'have a look at the inner thoughts of the trained models on these evals, and see whether they reference the constitution' field value experiment constitution references in reasoning traces, training corpus vs eval time (MASK, ODCV), for the four constitutional-SFT arms — Callum 2026-09-14: 'have a look at the inner thoughts… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-train-vs-eval-trait-ref.textn<1K0 likes49 downloads4d agoHugging Face24marin-dna /evals_traitgym_mendeliantabular1K<n<10K0 likes44 downloads8mo agoHugging Face25RuneForgeAI /Norse_Gods_and_Goddesses_Personality_Traits_Volume1 Norse Gods and Goddesses Personality Traits Volume 1 This dataset, titled Norse_Gods_and_Goddesses_Personality_Traits_Volume1, is a collection of synthetic conversational dialogues inspired by Norse mythology. It focuses on exploring the personality traits, myths, virtues, and philosophical insights of key Norse gods and goddesses such as Odin, Thor, Tyr, and Freyja. The dialogues are structured as role-played exchanges between fictional Norse characters (e.g., "Astrid" and "Bjorn")… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/Norse_Gods_and_Goddesses_Personality_Traits_Volume1.0 likes43 downloads9mo agoHugging Face26compass-group-tue /sdf_evaluation_traits_15M Models That Know How Evaluations Are Designed Score Safer This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.tabulartext-generation10K<n<100K0 likes42 downloads1mo agoHugging Face27curveball-steering /conversations_rude_llama3.2-3B-it-traits-v1_largetext10K<n<100K0 likes41 downloads12d agoHugging Face28morrislab /vep-traitgym-mrna Overview The variant effect prediction task measures the pathogenicity of single nucleotide polymorpism (SNPs). This dataset is a reprocessing of the TraitGym dataset (https://huggingface.co/datasets/songlab/TraitGym), see original dataset for data generation process. We have filtered TraitGym to only include SNPs in mature mRNA UTR regions, and provide the mRNA transcript sequence context for the SNP using the principle isoform as determined by APPRIS. This dataset is redistributed… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/vep-traitgym-mrna.text1K<n<10K0 likes38 downloads9mo agoHugging Face29curveball-steering /conversations_sadness_llama3.2-3B-it-traits-v1_largetext1K<n<10K0 likes38 downloads12d agoHugging Face30curveball-steering /conversations_excitement_llama3.2-3B-it-traits-v1_largetext1K<n<10K0 likes38 downloads12d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.