datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flowzap-sequence-workflows
sequence-workflows
A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case.
Organization Model
Top-level folders are primary Use Cases from the FlowZap Templates dropdown.
Second-level folders preserve the original source domain from the FlowZap app index.
Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes.
Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.openvid-frame-sequences-1M
OpenVid Frame Sequences — 1M adjacent frame pairs
Short, single-shot frame sequences cut from OpenVid-1M,
built to train and evaluate models on what changes between two frames half a second apart.
One sample = 10 consecutive frames, 0.5 s apart (a 4.5 s span) → 9 adjacent frame pairs.
[f00] --0.5s--> [f01] --0.5s--> [f02] ... [f09]
^ the thing you describe / predict
Sequences
116,596
Frames per sequence
10 (0.5 s apart, t = 0.0 … 4.5 s)
Adjacent frame… See the full description on the dataset page: https://huggingface.co/datasets/junha1125/openvid-frame-sequences-1M.omnimind-genomic-sequences-6sp
Genomic Sequences — 6 Species (20kb windows, Etapa C replication)
Sequências genômicas reais (NCBI/Ensembl) usadas no pipeline OmniMind Etapa C/D:
20.000 bp ACGT filtrados por espécie, mesmas regiões e janela da Etapa C
(nucleotide-transformer-v2-50m-multi-species).
Species
Chromosome
Source
Homo sapiens
22
NCBI
Saccharomyces cerevisiae
chrI
NCBI
Caenorhabditis elegans
chrI
NCBI
Drosophila melanogaster
2L
NCBI
Arabidopsis thaliana
1
NCBI
Escherichia coli… See the full description on the dataset page: https://huggingface.co/datasets/fabricioslv/omnimind-genomic-sequences-6sp.Sequence-of-action-prediction-mind2webhan-task-sequence-dataset-v1
Humanoid Task Sequence Dataset
This dataset contains example task sequences
used to train and validate humanoid task planning models.
It supports multi-step task execution
within the Humanoid Network.
Data Fields
Task list
Execution order
Task dependency
Format
JSON
Part of
Humanoid Network (HAN)
License
MIT
han-humanoid-intent-signal-sequences-v1
Humanoid Intent Signal Sequences
Overview
This dataset contains structured multimodal signal sequences
used for detecting human intent during interaction.
Signals include motion cues, speech tone indicators,
and proximity dynamics.
Data Fields
session_id
skeletal_motion_vector
speech_tone_vector
proximity_distance
timestamp_sequence
labeled_intent
Intended Use
Intent prediction modeling
Human-robot interaction research
Multimodal perception… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-intent-signal-sequences-v1.plate-washer-full-sequence-jul-30han-daily-assist-action-sequences-v1
Daily Assist Action Sequences
This dataset contains structured daily assistance
action sequences performed by humanoid robots
to support basic human activities.
The dataset focuses on practical, repeatable,
and safe household tasks.
Contents
Task name
Ordered action steps
Required tool
Completion state
Intended Use
Training sequence models for
household humanoid assistance systems.
License
MIT
adaption-ebolavirus-protein-sequences
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ebolavirus_protein_sequences
This dataset contains amino acid sequences for seven key proteins from various Ebola and Marburg virus genomes, including strains like Zaire, Sudan, and Tai Forest. Each entry provides the protein identifier, name, strain information, and the full sequence intended for generating embeddings using models like ESM-2 or ProtT5. The collection includes major… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-ebolavirus-protein-sequences.ebolavirus_protein_sequences_INITIAL
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ebolavirus_protein_sequences
This dataset contains amino acid sequences for seven key proteins from various Ebola and Marburg virus genomes, including strains like Zaire, Sudan, and Tai Forest. Each entry provides the protein identifier, name, strain information, and the full sequence intended for generating embeddings using models like ESM-2 or ProtT5. The collection includes major… See the full description on the dataset page: https://huggingface.co/datasets/joduor/ebolavirus_protein_sequences_INITIAL.code-with-ast-sequencecode-with-ast-sequence-comment-cleanedcode-with-ast-sequence-marked-indentmohiniyattam-pose-sequencescode-with-ast-sequence-comment-cleaned-hierarchicalcode-with-ast-sequence-hierarchical-2gsm8k_training_synthetic_positive_sequencehumanoid-action-sequence-datasetHumanoid Action Sequence Dataset
This dataset contains simple action sequences intended for humanoid
robots to understand ordered task execution and basic movement logic.
The data is designed for lightweight experimentation and benchmarking.
code-ast-node-edge-sequencecode-with-ast-sequence-hierarchicalsequence-results-sarvam-mswim_motion_sequences.jsonolmo2-7b-memorized-sequencescode-with-ast-sequence-comment-cleaned-hierarchical-annotatedsimple-task-sequence-datasetSimple Task Sequence Dataset
A dataset representing ordered steps for simple task execution.
obot_actions_sequence.jsonnumber-sequence-datasets
