CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01macwiatrak /bacbench-antibiotic-resistance-protein-sequences Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences) A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces separating different contigs present in the genome. The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-protein-sequences.text10K<n<100K0 likes703 downloads1y agoHugging Face02AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes694 downloads10d agoHugging Face03AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes659 downloads1mo agoHugging Face04bloyal /oas-paired-sequence-data Dataset Card for OAS Paired Sequence Data Dataset Summary Paired heavy- and light-chain sequence information from the Observed Antibody Space (OAS) database, downloaded on September 9, 2023. textfill-mask1M<n<10M1 likes537 downloads3y agoHugging Face05GenerTeam /sequence-recovery Next K-mer Prediction Abouts The Next K-mer Prediction task is a zero-shot evaluation method introduced in the GENERator paper to assess the quality of pretrained models. It involves inputting a sequence segment into the model and having it predict the next K base pairs. The predicted sequence is then compared to the actual sequence to assess accuracy. Sequence: The input sequence has a maximum length of 96k base pairs (bp). You can control the number of input… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/sequence-recovery.texttext-generation10K<n<100K8 likes475 downloads3mo agoHugging Face06AdoCleanCode /sequences_only_correct_V8text1M<n<10M0 likes442 downloads9mo agoHugging Face07mbafca2 /bacbench-antibiotic-resistance-protein-sequences Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences) A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces separating different contigs present in the genome. The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/mbafca2/bacbench-antibiotic-resistance-protein-sequences.text10K<n<100K0 likes424 downloads7mo agoHugging Face08tattabio /bac_16S_sequencestextn<1K0 likes351 downloads2y agoHugging Face09macwiatrak /bacbench-ppi-stringdb-protein-sequences Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.tabular1K<n<10K0 likes320 downloads1y agoHugging Face10another-phytophile /153-angiosperm-species-32k-sequences-shuffledtabular1M<n<10M0 likes317 downloads5mo agoHugging Face11viral-data-safety /sequence_homology_based_v2text10M<n<100M0 likes314 downloads4mo agoHugging Face12MA-tokenweights /pubmed-2019-pythia-word-tfidf-pubmedqa-clean-val-sequencestabular1K<n<10K0 likes279 downloads1mo agoHugging Face13MA-tokenweights /pubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-val-sequencestabular1K<n<10K0 likes277 downloads1mo agoHugging Face14cskokgibbs /yeast-gene-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes274 downloads1y agoHugging Face15macwiatrak /bacbench-phenotypic-traits-protein-sequences Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences) A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels. The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the bacterial genome, ordered by their location on the chromosome and plasmids. The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.text10K<n<100K0 likes239 downloads10mo agoHugging Face16Hiesh /robomme_sequencerecoveryvertically RoboMME — SequenceRecoveryVertically (Video QA) Video-QA dataset for the SequenceRecoveryVertically task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.tabularvideo-text-to-textn<1K0 likes185 downloads2mo agoHugging Face17Hiesh /robomme_sequencerecoveryhorizontally RoboMME — SequenceRecoveryHorizontally (Video QA) Video-QA dataset for the SequenceRecoveryHorizontally task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.tabularvideo-text-to-textn<1K0 likes147 downloads2mo agoHugging Face18yiyi159 /cityscapes_sequence_1024by512image100K<n<1M0 likes119 downloads1y agoHugging Face19neuralbioinfo /PhaStyle-SequenceDB Dataset Card for neuralbioinfo/PhaStyle-SequenceDB phastyle Sequence Database A collection of bacteriophage nucleotide sequences and metadata for training and evaluating phage lifestyle prediction models. Available splits support both strict-holdout and standard-holdout experiments. Dataset Features Name Type Description sequence_id int64 Unique integer identifier for each sequence dataset string Source collection name (see “Splits” below)… See the full description on the dataset page: https://huggingface.co/datasets/neuralbioinfo/PhaStyle-SequenceDB.tabular1K<n<10K0 likes99 downloads1y agoHugging Face20macwiatrak /bacbench-operon-identification-protein-sequences Dataset for operon identification in bacteria (Protein sequences) A dataset of 4,073 operons across 11 bacterial genomes species. The operon annotations have been extracted from Operon DB and the genome protein sequences have been extracted from GenBank. Each row contains a set of protein sequences present in the genome, represented by a list of protein sequences from different contigs. We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-protein-sequences.textn<1K0 likes98 downloads1y agoHugging Face21uclgroup8 /iemocap-embeddings-sequencestext1K<n<10K0 likes92 downloads3y agoHugging Face22hle2000 /Mintaka_Sequences_T5-xl-ssm Dataset Card for "Mintaka_Sequences_T5-xl-ssm" More Information needed text10K<n<100K0 likes88 downloads3y agoHugging Face23willdaspit /afdb_50_sequence_clustered_reprstabular10M<n<100M1 likes84 downloads11mo agoHugging Face24Aditya02 /Charades-Action-Sequence-Sampletabular1K<n<10K0 likes81 downloads2y agoHugging Face25cskokgibbs /yeast-tf-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes77 downloads1y agoHugging Face26macwiatrak /phenotypic-trait-catalase-protein-sequences Dataset for predicting Catalase phenotype from whole bacterial genomes (protein sequences) A dataset of over 1k bacterial genomes across species with the Catalase as label. Catalase denotes whether a bacterium produces the catalase enzyme that breaks down hydrogen peroxide (H₂O₂) into water and oxygen, thereby protecting the cell from oxidative stress. Here, we provide binary Catalase labels, therefore the problem is a binary classification problem. The genome protein sequences… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/phenotypic-trait-catalase-protein-sequences.text1K<n<10K0 likes76 downloads1y agoHugging Face27uclgroup8 /iemocap-embeddings-sequences-with-early-audiotext1K<n<10K0 likes74 downloads3y agoHugging Face28KomeijiForce /Fine_Grained_Fandom_Benchmark_Action_Sequences Codified Decision Tree (CDT) Action Sequences This dataset contains scene-action pairs derived from storylines, used to train and evaluate role-playing (RP) agents using the Codified Decision Trees (CDT) framework. Paper: Deriving Character Logic from Storyline as Codified Decision Trees Repository: https://github.com/KomeijiForce/Codified_Decision_Tree Introduction Role-playing (RP) agents rely on behavioral profiles to act consistently across diverse narrative… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Fine_Grained_Fandom_Benchmark_Action_Sequences.texttext-generation10K<n<100K0 likes72 downloads8mo agoHugging Face29tattabio /rpob_arch_dna_phylogeny_sequencestextn<1K0 likes71 downloads2y agoHugging Face30Atomi /XES3G5M_interaction_sequencestabular10K<n<100K0 likes71 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.