datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DyNativeGaussian_sequence
DyNativeGaussian Sequence
Demo: Free-Viewpoint Camera Move
Dataset Overview
DyNativeGaussian_sequence is a curated dynamic scene dataset for research on dynamic scene compression, dynamic novel view synthesis, 4D reconstruction, dynamic Gaussian Splatting, temporal rendering, and video-based scene representation learning.
The dataset contains multiple dynamic indoor, outdoor, and performance scenes, including VRU, N3DV, MeetRoom, and Dance-Dunhuang… See the full description on the dataset page: https://huggingface.co/datasets/LeeXiangNO1/DyNativeGaussian_sequence.flowzap-sequence-workflows
sequence-workflows
A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case.
Organization Model
Top-level folders are primary Use Cases from the FlowZap Templates dropdown.
Second-level folders preserve the original source domain from the FlowZap app index.
Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes.
Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.sequence_recovery_centered_200_horizontally
RoboMME — SequenceRecoveryHorizontally (Robot HDF5 Demonstrations)
Raw robot demonstration data for the SequenceRecoveryHorizontally task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. Each HDF5 file is one recorded episode
containing observations (RGB/state), actions, and metadata for imitation
learning.
Layout
train/ 100 episodes
val/ 50 episodes
test/ 50 episodes
Files are named episode_<idx>_seed_<seed>.h5.… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/sequence_recovery_centered_200_horizontally.sequence_recovery_centered_200_vertically
RoboMME — SequenceRecoveryVertically (Robot HDF5 Demonstrations)
Raw robot demonstration data for the SequenceRecoveryVertically task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. Each HDF5 file is one recorded episode
containing observations (RGB/state), actions, and metadata for imitation
learning.
Layout
train/ 100 episodes
val/ 50 episodes
test/ 50 episodes
Files are named episode_<idx>_seed_<seed>.h5. Seeds… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/sequence_recovery_centered_200_vertically.bacbench-antibiotic-resistance-protein-sequences
Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences)
A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces
separating different contigs present in the genome.
The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-protein-sequences.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.carbon-cpu-enriched-sequences-sampledoas-paired-sequence-data
Dataset Card for OAS Paired Sequence Data
Dataset Summary
Paired heavy- and light-chain sequence information from the Observed Antibody Space (OAS) database, downloaded on September 9, 2023.
sequence-recovery
Next K-mer Prediction
Abouts
The Next K-mer Prediction task is a zero-shot evaluation method introduced in the GENERator paper to assess the quality of pretrained models. It involves inputting a sequence segment into the model and having it predict the next K base pairs. The predicted sequence is then compared to the actual sequence to assess accuracy.
Sequence: The input sequence has a maximum length of 96k base pairs (bp). You can control the number of input… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/sequence-recovery.bacbench-antibiotic-resistance-protein-sequences
Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences)
A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces
separating different contigs present in the genome.
The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/mbafca2/bacbench-antibiotic-resistance-protein-sequences.sabdab_joint_sequences_uniprothandball_video_sequencessequences_only_correct_V8bac_16S_sequencesbacbench-phenotypic-traits-protein-sequences
Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the
bacterial genome, ordered by their location on the chromosome and plasmids.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.robomme_sequencerecoveryvertically
RoboMME — SequenceRecoveryVertically (Video QA)
Video-QA dataset for the SequenceRecoveryVertically task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.sequence_homology_based_v2robomme_sequencerecoveryhorizontally
RoboMME — SequenceRecoveryHorizontally (Video QA)
Video-QA dataset for the SequenceRecoveryHorizontally task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.153-angiosperm-species-32k-sequences-shuffledpubmed-2019-pythia-word-tfidf-pubmedqa-clean-val-sequencesflip2-multi-sequence-prompt-ablation-generated-variants-2pubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-val-sequencesbacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.Origin-Sequence-Data
AL-GR/Origin-Sequence-Data: Raw User Behavior Sequences 📜
About the Dataset
Each row in this dataset (Origin-Sequence-Data) represents a step in a user's journey, consisting of a sequence of previously interacted items (user_history) and the next item they interacted with (target_item). All item IDs have been anonymized into short, unique strings.
This dataset is ideal for:
🧑🔬 Researchers who want to design their own data processing or prompting strategies for… See the full description on the dataset page: https://huggingface.co/datasets/AL-GR/Origin-Sequence-Data.yeast-gene-sequence-homology-pretokenized-NTall-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequencesall-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequencesdaily-paper-2026-08-31-multi-skill-gap-sequence-routing
The Multi-Skill Gap: Measuring the Cost-Quality Frontier of Order-Sensitive Skill-Sequence Routing in a 2,200-Skill Agent Harness
TL;DR — First order-sensitive chain benchmark for skill routing: 48 composite tasks (12 real workflow templates x 4 variants) over a 2,275-skill bilingual production registry, measured across free BM25 top-k retrieval, self-hosted Qwen3-1.7B/14B composers, and a frontier cost anchor. Free top-20 retrieval reaches 75.7% set completeness but 0.0% exact… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-08-31-multi-skill-gap-sequence-routing.plant-promoter-sequences
Promoter Sequences for Various plant species
The data in this dataset has the promoter sequences for 241 different plant species and has been used for the pretraining step of Florabert. It has been created by processing the raw fasta files and the gff3 / gff files from Ensembl and Refseq.
samtools and bedtools have been used to extract the promoter sequences from these that are 1kb upstream of the sequence.
The data has been split into train and test data (90-10 split). In all… See the full description on the dataset page: https://huggingface.co/datasets/Gurveer05/plant-promoter-sequences.nonmonotonic_sequence_generation_checkpoints
