CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Jules-OC /flowzap-sequence-workflows sequence-workflows A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case. Organization Model Top-level folders are primary Use Cases from the FlowZap Templates dropdown. Second-level folders preserve the original source domain from the FlowZap app index. Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes. Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.tabularn<1K0 likes1.2k downloads6mo agoHugging Face02AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes729 downloads7d agoHugging Face03AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes649 downloads27d agoHugging Face04Hiesh /robomme_sequencerecoveryvertically RoboMME — SequenceRecoveryVertically (Video QA) Video-QA dataset for the SequenceRecoveryVertically task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.tabularvideo-text-to-textn<1K0 likes324 downloads2mo agoHugging Face05Hiesh /robomme_sequencerecoveryhorizontally RoboMME — SequenceRecoveryHorizontally (Video QA) Video-QA dataset for the SequenceRecoveryHorizontally task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.tabularvideo-text-to-textn<1K0 likes310 downloads2mo agoHugging Face06another-phytophile /153-angiosperm-species-32k-sequences-shuffledtabular1M<n<10M0 likes309 downloads5mo agoHugging Face07MA-tokenweights /pubmed-2019-pythia-word-tfidf-pubmedqa-clean-val-sequencestabular1K<n<10K0 likes304 downloads27d agoHugging Face08MA-tokenweights /pubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-val-sequencestabular1K<n<10K0 likes277 downloads26d agoHugging Face09macwiatrak /bacbench-ppi-stringdb-protein-sequences Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.tabular1K<n<10K0 likes261 downloads1y agoHugging Face10cskokgibbs /yeast-gene-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes216 downloads1y agoHugging Face11MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes201 downloads29d agoHugging Face12MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes174 downloads29d agoHugging Face13MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes112 downloads1mo agoHugging Face14MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes88 downloads1mo agoHugging Face15willdaspit /afdb_50_sequence_clustered_reprstabular10M<n<100M1 likes83 downloads11mo agoHugging Face16neuralbioinfo /PhaStyle-SequenceDB Dataset Card for neuralbioinfo/PhaStyle-SequenceDB phastyle Sequence Database A collection of bacteriophage nucleotide sequences and metadata for training and evaluating phage lifestyle prediction models. Available splits support both strict-holdout and standard-holdout experiments. Dataset Features Name Type Description sequence_id int64 Unique integer identifier for each sequence dataset string Source collection name (see “Splits” below)… See the full description on the dataset page: https://huggingface.co/datasets/neuralbioinfo/PhaStyle-SequenceDB.tabular1K<n<10K0 likes80 downloads1y agoHugging Face17Aditya02 /Charades-Action-Sequence-Sampletabular1K<n<10K0 likes79 downloads2y agoHugging Face18Atomi /XES3G5M_interaction_sequencestabular10K<n<100K0 likes78 downloads2y agoHugging Face19junha1125 /openvid-frame-sequences-1M OpenVid Frame Sequences — 1M adjacent frame pairs Short, single-shot frame sequences cut from OpenVid-1M, built to train and evaluate models on what changes between two frames half a second apart. One sample = 10 consecutive frames, 0.5 s apart (a 4.5 s span) → 9 adjacent frame pairs. [f00] --0.5s--> [f01] --0.5s--> [f02] ... [f09] ^ the thing you describe / predict Sequences 116,596 Frames per sequence 10 (0.5 s apart, t = 0.0 … 4.5 s) Adjacent frame… See the full description on the dataset page: https://huggingface.co/datasets/junha1125/openvid-frame-sequences-1M.tabularimage-to-text100K<n<1M0 likes74 downloads1mo agoHugging Face20cskokgibbs /yeast-tf-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes72 downloads1y agoHugging Face21random-sequence /flock-demo-critical-infra-sectionstabular10K<n<100K0 likes71 downloads7mo agoHugging Face22willdaspit /afdb_50_sequence_clustered_reprs_GOGO functional annotations are semicolon-separated in go_ids. The "GO:" prefix is stripped. Only ids appearing at least 10000 times in the training set are kept - there are 425 such ids. Parents are automatically populated (e.g. iron binding -> metal binding). Obsolete ids are replaced with current where possible, or removed if not. tabular10M<n<100M0 likes62 downloads7mo agoHugging Face23CodeIsAbstract /sanskrit-morpho-sequences Sanskrit Morphological Sequence Corpus (Vidyut-Verified) A large-scale, Pāṇinian-verified morphological sequence dataset for classical and Vedic Sanskrit. Every token is annotated with its lemma, generative root (aupadeśika), part-of-speech, case, number, person, voice, and gender — all in the SLP1 transliteration, and all aligned at the sentence level for sequence-tagging / seq2seq training. 710,785 sentences (after deduplication) 5,511,664 tokens 14 columns (10 linguistic + 4… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-morpho-sequences.tabulartoken-classification100K<n<1M0 likes58 downloads2mo agoHugging Face24ClarusC64 /quantum-gate-sequence-instability-v0.1 quantum-gate-sequence-instability-v0.1 What this dataset does This dataset evaluates whether models can detect instability in quantum gate sequences. Each row represents a simplified quantum circuit execution scenario described through observable device and circuit proxies. The task is to determine whether the gate sequence remains executable inside a stable coherence window or becomes unstable. Core stability idea Quantum gate sequences become unstable when… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-gate-sequence-instability-v0.1.tabulartabular-classificationn<1K0 likes47 downloads5mo agoHugging Face25JasonTWalker /tokenized_uniprotkb_1024_sequence_lengthtabular100K<n<1M0 likes40 downloads1y agoHugging Face26willdaspit /afdb_50_sequence_clusteredAnnotated sequences from the AFDB50 (sequence-based, not structure-based) clustering at https://afdb-cluster.steineggerlab.workers.dev/ Two rows with the same RepId are part of the same cluster. Similarly, two rows with the same cluster_id are part of the same cluster. cluster_id is in order from smallest to largest cluster - that is, members of the smallest cluster have cluster_id=0. Singletons are included. All plddts are included. All are annotated with the number of cluster members, plddt… See the full description on the dataset page: https://huggingface.co/datasets/willdaspit/afdb_50_sequence_clustered.tabular100M<n<1B1 likes37 downloads11mo agoHugging Face27XingweiT /IntrEx-sequence IntrEx: A Dataset for Modeling Engagement in Educational Conversations (sequence-level) 【 📦 GitHub repo | 🤗 Paper 】 TL;DR IntrEx is the first large-scale dataset annotated for interestingness and expected interestingness in teacher-student interactions. Data Fields Column Description project_id ID for specifying a unit of annotation work where a batch of participants annotate a set of conversations page_id The annotation page number inside… See the full description on the dataset page: https://huggingface.co/datasets/XingweiT/IntrEx-sequence.tabulartext-classification1K<n<10K1 likes31 downloads1y agoHugging Face28another-phytophile /153-angiosperm-species-8192bp-sequencestabular1M<n<10M0 likes30 downloads6mo agoHugging Face29herutriana44 /Drugbank_Summary_Drug_Sequencetabularn<1K1 likes25 downloads3y agoHugging Face30oscarz511 /arm_sequencestabular1K<n<10K0 likes25 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.