datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flowzap-sequence-workflows
sequence-workflows
A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case.
Organization Model
Top-level folders are primary Use Cases from the FlowZap Templates dropdown.
Second-level folders preserve the original source domain from the FlowZap app index.
Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes.
Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.carbon-cpu-enriched-sequences-sampledrobomme_sequencerecoveryvertically
RoboMME — SequenceRecoveryVertically (Video QA)
Video-QA dataset for the SequenceRecoveryVertically task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.robomme_sequencerecoveryhorizontally
RoboMME — SequenceRecoveryHorizontally (Video QA)
Video-QA dataset for the SequenceRecoveryHorizontally task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.153-angiosperm-species-32k-sequences-shuffledpubmed-2019-pythia-word-tfidf-pubmedqa-clean-val-sequencespubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-val-sequencesbacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.yeast-gene-sequence-homology-pretokenized-NTall-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequencesall-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequenceswikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-val-sequenceswikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-val-sequencesafdb_50_sequence_clustered_reprsPhaStyle-SequenceDB
Dataset Card for neuralbioinfo/PhaStyle-SequenceDB
phastyle Sequence Database
A collection of bacteriophage nucleotide sequences and metadata for training and evaluating phage lifestyle prediction models. Available splits support both strict-holdout and standard-holdout experiments.
Dataset Features
Name
Type
Description
sequence_id
int64
Unique integer identifier for each sequence
dataset
string
Source collection name (see “Splits” below)… See the full description on the dataset page: https://huggingface.co/datasets/neuralbioinfo/PhaStyle-SequenceDB.Charades-Action-Sequence-SampleXES3G5M_interaction_sequencesopenvid-frame-sequences-1M
OpenVid Frame Sequences — 1M adjacent frame pairs
Short, single-shot frame sequences cut from OpenVid-1M,
built to train and evaluate models on what changes between two frames half a second apart.
One sample = 10 consecutive frames, 0.5 s apart (a 4.5 s span) → 9 adjacent frame pairs.
[f00] --0.5s--> [f01] --0.5s--> [f02] ... [f09]
^ the thing you describe / predict
Sequences
116,596
Frames per sequence
10 (0.5 s apart, t = 0.0 … 4.5 s)
Adjacent frame… See the full description on the dataset page: https://huggingface.co/datasets/junha1125/openvid-frame-sequences-1M.yeast-tf-sequence-homology-pretokenized-NTflock-demo-critical-infra-sectionsafdb_50_sequence_clustered_reprs_GOGO functional annotations are semicolon-separated in go_ids.
The "GO:" prefix is stripped.
Only ids appearing at least 10000 times in the training set are kept - there are 425 such ids.
Parents are automatically populated (e.g. iron binding -> metal binding). Obsolete ids are replaced with current where possible, or removed if not.
sanskrit-morpho-sequences
Sanskrit Morphological Sequence Corpus (Vidyut-Verified)
A large-scale, Pāṇinian-verified morphological sequence dataset for
classical and Vedic Sanskrit. Every token is annotated with its lemma,
generative root (aupadeśika), part-of-speech, case, number, person, voice,
and gender — all in the SLP1 transliteration, and all aligned at the
sentence level for sequence-tagging / seq2seq training.
710,785 sentences (after deduplication)
5,511,664 tokens
14 columns (10 linguistic + 4… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-morpho-sequences.quantum-gate-sequence-instability-v0.1
quantum-gate-sequence-instability-v0.1
What this dataset does
This dataset evaluates whether models can detect instability in quantum gate sequences.
Each row represents a simplified quantum circuit execution scenario described through observable device and circuit proxies.
The task is to determine whether the gate sequence remains executable inside a stable coherence window or becomes unstable.
Core stability idea
Quantum gate sequences become unstable when… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-gate-sequence-instability-v0.1.tokenized_uniprotkb_1024_sequence_lengthafdb_50_sequence_clusteredAnnotated sequences from the AFDB50 (sequence-based, not structure-based) clustering at https://afdb-cluster.steineggerlab.workers.dev/
Two rows with the same RepId are part of the same cluster. Similarly, two rows with the same cluster_id are part of the same cluster.
cluster_id is in order from smallest to largest cluster - that is, members of the smallest cluster have cluster_id=0.
Singletons are included.
All plddts are included.
All are annotated with the number of cluster members, plddt… See the full description on the dataset page: https://huggingface.co/datasets/willdaspit/afdb_50_sequence_clustered.IntrEx-sequence
IntrEx: A Dataset for Modeling Engagement in Educational Conversations (sequence-level)
【 📦 GitHub repo | 🤗 Paper 】
TL;DR
IntrEx is the first large-scale dataset annotated for interestingness and expected interestingness in teacher-student interactions.
Data Fields
Column
Description
project_id
ID for specifying a unit of annotation work where a batch of participants annotate a set of conversations
page_id
The annotation page number inside… See the full description on the dataset page: https://huggingface.co/datasets/XingweiT/IntrEx-sequence.153-angiosperm-species-8192bp-sequencesDrugbank_Summary_Drug_Sequencearm_sequences
