datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodonTranslator-data
CodonTranslator Data
This repository contains the final public training-data release used for CodonTranslator.
Contents
train/: representative-only training shards
val/: representative-only validation shards
test/: representative-only held-out test shards
embeddings_v2/: precomputed species conditioning embeddings used in training
_work/final_representative_counts.json: final released split sizes
_work/split_report.json: split audit report… See the full description on the dataset page: https://huggingface.co/datasets/alegendaryfish/CodonTranslator-data.CodonTransformer
CodonTransformer Dataset
A comprehensive compilation of 1,001,197 DNA and protein sequence pairs, sourced from 164 organisms across Eukaryotes, Bacteria, and Archaea.
This dataset provides a rich resource for various computational biology and bioinformatics applications such as studying gene sequences, codon usage, and protein expression across diverse species.
Dataset Contents
1,001,197 DNA-protein sequence pairs
Sequences from 164 organisms, including:
Eukaryotes:… See the full description on the dataset page: https://huggingface.co/datasets/adibvafa/CodonTransformer.codons
Fungal coding sequence dataset
Dataset of codon usage for fungal organisms created from the Ensembl Genomes clustered to 50% sequence identity at the protein level and split into 80%/10%/10% train/validation/test splits for use in training a neural network to design native-looking nucleotide sequences for fungal organisms
Dataset processing
This document describes the preparation of the fungal codons
dataset.
Obtaining the raw data
The raw data, CDS sequences… See the full description on the dataset page: https://huggingface.co/datasets/alxcarln/codons.Motif-Prev1
