datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.carbon-cpu-enriched-sequences-sampled153-angiosperm-species-32k-sequences-shuffledpubmed-2019-pythia-word-tfidf-pubmedqa-clean-val-sequencespubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-val-sequencesbacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.all-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequencesall-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequenceswikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-val-sequenceswikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-val-sequencesXES3G5M_interaction_sequencesopenvid-frame-sequences-1M
OpenVid Frame Sequences — 1M adjacent frame pairs
Short, single-shot frame sequences cut from OpenVid-1M,
built to train and evaluate models on what changes between two frames half a second apart.
One sample = 10 consecutive frames, 0.5 s apart (a 4.5 s span) → 9 adjacent frame pairs.
[f00] --0.5s--> [f01] --0.5s--> [f02] ... [f09]
^ the thing you describe / predict
Sequences
116,596
Frames per sequence
10 (0.5 s apart, t = 0.0 … 4.5 s)
Adjacent frame… See the full description on the dataset page: https://huggingface.co/datasets/junha1125/openvid-frame-sequences-1M.sanskrit-morpho-sequences
Sanskrit Morphological Sequence Corpus (Vidyut-Verified)
A large-scale, Pāṇinian-verified morphological sequence dataset for
classical and Vedic Sanskrit. Every token is annotated with its lemma,
generative root (aupadeśika), part-of-speech, case, number, person, voice,
and gender — all in the SLP1 transliteration, and all aligned at the
sentence level for sequence-tagging / seq2seq training.
710,785 sentences (after deduplication)
5,511,664 tokens
14 columns (10 linguistic + 4… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-morpho-sequences.153-angiosperm-species-8192bp-sequencesarm_sequences153-angiosperm-species-8192bp-sequences-balanced-for-autointerphuman_genome_gnomAD_doped_sequences_v3wikitext-103-raw-v2-tfidf-invfreq-topic-stratified-v1-val-sequenceswikitext-103-raw-pythia-tfidf-topic-stratified-v1-val-sequenceswikitext-103-raw-pythia-tfidf-invfreq-topic-stratified-v1-val-sequencesall-the-news-2-tfidf-invfreq-topic-stratified-v1-val-sequencespubmed-tfidf-linear-token-val-sequencessteered_sequences_3B_multiplepubmed-tfidf-sublinear-token-val-sequenceswikitext-103-raw-v2-tfidf-topic-stratified-v1-val-sequences-sublinearall-the-news-2-freq-invsqrt-topic-stratified-v1-val-sequenceswikitext-103-raw-v2-tfidf-topic-stratified-v1-val-sequencesall-the-news-2-tfidf-topic-stratified-v1-val-sequencesall-the-news-2-tfidf-topic-stratified-v1-val-sequences-wordalignedall-the-news-2-tfidf-topic-stratified-v1-val-sequences-sublinear
