datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flowzap-sequence-workflows
sequence-workflows
A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case.
Organization Model
Top-level folders are primary Use Cases from the FlowZap Templates dropdown.
Second-level folders preserve the original source domain from the FlowZap app index.
Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes.
Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.carbon-cpu-enriched-sequences-sampledbacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.153-angiosperm-species-32k-sequences-shuffledpubmed-2019-pythia-word-tfidf-pubmedqa-clean-val-sequencespubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-val-sequencesyeast-gene-sequence-homology-pretokenized-NTrobomme_sequencerecoveryvertically
RoboMME — SequenceRecoveryVertically (Video QA)
Video-QA dataset for the SequenceRecoveryVertically task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.robomme_sequencerecoveryhorizontally
RoboMME — SequenceRecoveryHorizontally (Video QA)
Video-QA dataset for the SequenceRecoveryHorizontally task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.PhaStyle-SequenceDB
Dataset Card for neuralbioinfo/PhaStyle-SequenceDB
phastyle Sequence Database
A collection of bacteriophage nucleotide sequences and metadata for training and evaluating phage lifestyle prediction models. Available splits support both strict-holdout and standard-holdout experiments.
Dataset Features
Name
Type
Description
sequence_id
int64
Unique integer identifier for each sequence
dataset
string
Source collection name (see “Splits” below)… See the full description on the dataset page: https://huggingface.co/datasets/neuralbioinfo/PhaStyle-SequenceDB.afdb_50_sequence_clustered_reprsCharades-Action-Sequence-Sampleyeast-tf-sequence-homology-pretokenized-NTall-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequencesXES3G5M_interaction_sequencesflock-demo-critical-infra-sectionsopenvid-frame-sequences-1M
OpenVid Frame Sequences — 1M adjacent frame pairs
Short, single-shot frame sequences cut from OpenVid-1M,
built to train and evaluate models on what changes between two frames half a second apart.
One sample = 10 consecutive frames, 0.5 s apart (a 4.5 s span) → 9 adjacent frame pairs.
[f00] --0.5s--> [f01] --0.5s--> [f02] ... [f09]
^ the thing you describe / predict
Sequences
116,596
Frames per sequence
10 (0.5 s apart, t = 0.0 … 4.5 s)
Adjacent frame… See the full description on the dataset page: https://huggingface.co/datasets/junha1125/openvid-frame-sequences-1M.afdb_50_sequence_clustered_reprs_GOGO functional annotations are semicolon-separated in go_ids.
The "GO:" prefix is stripped.
Only ids appearing at least 10000 times in the training set are kept - there are 425 such ids.
Parents are automatically populated (e.g. iron binding -> metal binding). Obsolete ids are replaced with current where possible, or removed if not.
sanskrit-morpho-sequences
Sanskrit Morphological Sequence Corpus (Vidyut-Verified)
A large-scale, Pāṇinian-verified morphological sequence dataset for
classical and Vedic Sanskrit. Every token is annotated with its lemma,
generative root (aupadeśika), part-of-speech, case, number, person, voice,
and gender — all in the SLP1 transliteration, and all aligned at the
sentence level for sequence-tagging / seq2seq training.
710,785 sentences (after deduplication)
5,511,664 tokens
14 columns (10 linguistic + 4… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-morpho-sequences.all-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequencesopenarm-restock-sequences-canonical-30fps-subtasks-gripper-vlm
restock-sequences-canonical-30fps
LeRobot v2.1 dataset: 226 episodes, 306218 frames at 30 fps.
Robot: openarm_bimanual
Cameras: observation.images.context, observation.images.wrist_left, observation.images.wrist_right
State/action dim: 16
Load it with the v2.1 tag, which is the revision the training path pins.
quantum-gate-sequence-instability-v0.1
quantum-gate-sequence-instability-v0.1
What this dataset does
This dataset evaluates whether models can detect instability in quantum gate sequences.
Each row represents a simplified quantum circuit execution scenario described through observable device and circuit proxies.
The task is to determine whether the gate sequence remains executable inside a stable coherence window or becomes unstable.
Core stability idea
Quantum gate sequences become unstable when… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-gate-sequence-instability-v0.1.tokenized_uniprotkb_1024_sequence_lengthafdb_50_sequence_clusteredAnnotated sequences from the AFDB50 (sequence-based, not structure-based) clustering at https://afdb-cluster.steineggerlab.workers.dev/
Two rows with the same RepId are part of the same cluster. Similarly, two rows with the same cluster_id are part of the same cluster.
cluster_id is in order from smallest to largest cluster - that is, members of the smallest cluster have cluster_id=0.
Singletons are included.
All plddts are included.
All are annotated with the number of cluster members, plddt… See the full description on the dataset page: https://huggingface.co/datasets/willdaspit/afdb_50_sequence_clustered.IntrEx-sequence
IntrEx: A Dataset for Modeling Engagement in Educational Conversations (sequence-level)
【 📦 GitHub repo | 🤗 Paper 】
TL;DR
IntrEx is the first large-scale dataset annotated for interestingness and expected interestingness in teacher-student interactions.
Data Fields
Column
Description
project_id
ID for specifying a unit of annotation work where a batch of participants annotate a set of conversations
page_id
The annotation page number inside… See the full description on the dataset page: https://huggingface.co/datasets/XingweiT/IntrEx-sequence.153-angiosperm-species-8192bp-sequences153-angiosperm-species-8192bp-sequences-balanced-for-autointerpsoliaudit-dasp-sequence-gnn-no-explainerclinical-control-sequence-sepsis-v1
Clinical Control Sequence Sepsis Detection
Overview
This dataset tests whether a model can detect whether a proposed intervention sequence is the correct stabilizing control sequence for a sepsis-like clinical system.
Complex systems are often not stabilized by a single action. They require the correct sequence of interventions delivered in the correct order as the system evolves.
The goal of this benchmark is to determine whether the control sequence meaningfully guides… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-control-sequence-sepsis-v1.
