datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nucleotide_transformer_downstream_tasks
Dataset Card for Dataset Name
The nucleotide_transformer_downstream_tasks dataset features the 18 downstream tasks presented in the Nucleotide Transformer paper. They consist of both binary and multi-class classification tasks that aim at providing a consistent genomics benchmark.
⚠️We note that we have revised and improved our benchmark during the peer-review process. The datasets featured in this repository are used up to this release. We highly encourage to move to the new… See the full description on the dataset page: https://huggingface.co/datasets/InstaDeepAI/nucleotide_transformer_downstream_tasks.nucleotide_transformer_downstream_tasks_revised
Dataset Card for Dataset Name
The nucleotide_transformer_downstream_tasks dataset features the 18 downstream tasks presented in the Nucleotide Transformer paper. They consist of both binary and multi-class classification tasks that aim at providing a consistent genomics benchmark.
We note that this is an updated version of this benchmark after the paper has been through peer-review. We highly encourage to move to this version in detriment of the older version.Keypoints about the… See the full description on the dataset page: https://huggingface.co/datasets/InstaDeepAI/nucleotide_transformer_downstream_tasks_revised.cell-downstream-tasksversion https://git-lfs.github.com/spec/v1
oid sha256:6ab2d609b032deb880228678b4fc25f7c3c5edc33502a453ceabe41d6e0ae6b0
size 3033
rna-downstream-tasks
GB.RNA Benchmark Datasets
mRNA related tasks
Translation efficiency prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
mRNA expression level prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
Mean ribosome load prediction from Sample et al. (2019) [2]
input sequence: 5'UTR
ouput: mean ribosome load
the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.LucaOne-Downstream-Task-Datasets
split: devpath: dataset/ncRNAFam/gene/multi_class/dev/dev.csv
split: testpath: dataset/ncRNAFam/gene/multi_class/test/test.csv
split: labelpath: dataset/ncRNAFam/gene/multi_class/label.txt
config_name: ncRPIdata_files:
split: trainpath: dataset/ncRPI/gene_protein/binary_class/train/train.csv
split: devpath: dataset/ncRPI/gene_protein/binary_class/dev/dev.csv
split: testpath:… See the full description on the dataset page: https://huggingface.co/datasets/LucaGroup/LucaOne-Downstream-Task-Datasets.tissue-downstream-tasks
GB.Tissue Dataset Collection
niche type classification
Niche is the microenvironment in which each cell exists and is able to keep its own peculiar characteristics (Giacomo Donati, 2015). Based on spatial transcriptomic data, one can annotate niche label with established tools, which integrates similarity in gene expression profiles, spatial neighborhood structures and histological information in the tissue.
The task is to predict niche type of each cell given… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/tissue-downstream-tasks.toy_downstream_tasks_multilabelMultilabel datasets used in the Nucleotide Transformer paper.LLM_Description_Vocab_opt_facebook_opt_30b_downstream_tasks
Dataset Card for "LLM_Description_Vocab_opt_facebook_opt_30b_downstream_tasks"
More Information needed
LLM_Description_Vocab_opt_Multimodal_Fatima_opt_175b_downstream_tasks
Dataset Card for "LLM_Description_Vocab_opt_Multimodal_Fatima_opt_175b_downstream_tasks"
More Information needed
LLM_Description_Vocab_bloom_bigscience_bloom_downstream_tasks
Dataset Card for "LLM_Description_Vocab_bloom_bigscience_bloom_downstream_tasks"
More Information needed
TATA-NOTATA-FineMistral-nucleotide_transformer_downstream_tasksDataset for Fine-tuning Mistral Model on Tata and No Tata Sequences
Description
This dataset is specifically curated for training the Mistral model to distinguish between 'tata' and 'no tata' sequences. It is derived and reformatted from a dataset originally created by InstaDeep, tailored to enhance the performance of natural language processing models in identifying specific patterns.
Dataset Information
Features: This dataset consists of sequences represented as strings under the… See the full description on the dataset page: https://huggingface.co/datasets/Kamka-IT/TATA-NOTATA-FineMistral-nucleotide_transformer_downstream_tasks.LLM_Description_Vocab_gpt_3_text_davinci_003_downstream_tasksOxfordPets_facebook_opt_30b_LLM_Description_opt30b_downstream_tasks_ViT_L_14
Dataset Card for "OxfordPets_facebook_opt_30b_LLM_Description_opt30b_downstream_tasks_ViT_L_14"
More Information needed
OxfordPets_Multimodal_Fatima_opt_175b_LLM_Description_opt175b_downstream_tasks_ViT_L_14
Dataset Card for "OxfordPets_Multimodal_Fatima_opt_175b_LLM_Description_opt175b_downstream_tasks_ViT_L_14"
More Information needed
OxfordPets_facebook_opt_350m_LLM_Description_gpt3_downstream_tasks_ViT_L_14
Dataset Card for "OxfordPets_facebook_opt_350m_LLM_Description_gpt3_downstream_tasks_ViT_L_14"
More Information needed
downstream_tasks
