datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodonTransformer
CodonTransformer Dataset
A comprehensive compilation of 1,001,197 DNA and protein sequence pairs, sourced from 164 organisms across Eukaryotes, Bacteria, and Archaea.
This dataset provides a rich resource for various computational biology and bioinformatics applications such as studying gene sequences, codon usage, and protein expression across diverse species.
Dataset Contents
1,001,197 DNA-protein sequence pairs
Sequences from 164 organisms, including:
Eukaryotes:… See the full description on the dataset page: https://huggingface.co/datasets/adibvafa/CodonTransformer.codons
Fungal coding sequence dataset
Dataset of codon usage for fungal organisms created from the Ensembl Genomes clustered to 50% sequence identity at the protein level and split into 80%/10%/10% train/validation/test splits for use in training a neural network to design native-looking nucleotide sequences for fungal organisms
Dataset processing
This document describes the preparation of the fungal codons
dataset.
Obtaining the raw data
The raw data, CDS sequences… See the full description on the dataset page: https://huggingface.co/datasets/alxcarln/codons.
