datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GO-MO
GO-MO: A large-scale graph-augmented traffic dataset for data-driven spatio-temporal traffic analysis
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/GO-MO.go-mo-dataset
GO-MO, a massive Graph agumented Open urban MObility dataset
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).
Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.GO_MF
GO-MF Dataset
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF.GO_MF_ESMFold
GO-MF Dataset with ESMFold Structural Sequence
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF_ESMFold.GO_MF_AlphaFold2
GO-MF Dataset with AlphaFold2 Structural Sequence
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF_AlphaFold2.
