datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodonTranslator-data
CodonTranslator Data
This repository contains the final public training-data release used for CodonTranslator.
Contents
train/: representative-only training shards
val/: representative-only validation shards
test/: representative-only held-out test shards
embeddings_v2/: precomputed species conditioning embeddings used in training
_work/final_representative_counts.json: final released split sizes
_work/split_report.json: split audit report… See the full description on the dataset page: https://huggingface.co/datasets/alegendaryfish/CodonTranslator-data.codon-transformer-assets
LevinHarness/codon-transformer-assets — internal asset cache
This private dataset is an internal staging cache of third-party runtime
assets required by the Levin Harness plugin(s) listed below. It is not an
official distribution: nothing here is published under this account's own
terms, and it is not affiliated with or endorsed by any upstream project.
Ownership and licensing
Every file remains the property of its upstream authors.
Each file keeps its upstream… See the full description on the dataset page: https://huggingface.co/datasets/LevinHarness/codon-transformer-assets.Organic-Reasoning-195k
Organic-CoT: High-Fidelity Organic Chain-of-Thought & Deep Logical Reasoning Dataset
Reject mechanical replies. Replicate expert-level "Organic Thinking" and high-fidelity knowledge.
Deeply inspired by the Google Gemini Chain-of-Thought format, this dataset uses bilingual, high-precision synthetic SFT data to teach models how to perform "trial-and-error, reflection, and strategic planning" like human experts, rather than outputting rigid lists of steps.
Does Size… See the full description on the dataset page: https://huggingface.co/datasets/CodonProject/Organic-Reasoning-195k.CodonTransformer
CodonTransformer Dataset
A comprehensive compilation of 1,001,197 DNA and protein sequence pairs, sourced from 164 organisms across Eukaryotes, Bacteria, and Archaea.
This dataset provides a rich resource for various computational biology and bioinformatics applications such as studying gene sequences, codon usage, and protein expression across diverse species.
Dataset Contents
1,001,197 DNA-protein sequence pairs
Sequences from 164 organisms, including:
Eukaryotes:… See the full description on the dataset page: https://huggingface.co/datasets/adibvafa/CodonTransformer.codon-transformer-assets
LevinHarness/codon-transformer-assets — internal asset cache
This private dataset is an internal staging cache of third-party runtime
assets required by the Levin Harness plugin(s) listed below. It is not an
official distribution: nothing here is published under this account's own
terms, and it is not affiliated with or endorsed by any upstream project.
Ownership and licensing
Every file remains the property of its upstream authors.
Each file keeps its upstream… See the full description on the dataset page: https://huggingface.co/datasets/sgetttt/codon-transformer-assets.codons
Fungal coding sequence dataset
Dataset of codon usage for fungal organisms created from the Ensembl Genomes clustered to 50% sequence identity at the protein level and split into 80%/10%/10% train/validation/test splits for use in training a neural network to design native-looking nucleotide sequences for fungal organisms
Dataset processing
This document describes the preparation of the fungal codons
dataset.
Obtaining the raw data
The raw data, CDS sequences… See the full description on the dataset page: https://huggingface.co/datasets/alxcarln/codons.codontransformerMotif-Prev1genomiratheon_benchmark
GENOMIRATHEON™ Benchmark dataset
This repository contains the official GENOMIRATHEON™ prompt-response dataset, designed for evaluating and training language models on synthetic biology compliance simulations.
Dataset Overview:
genomiratheon_benchmark.json is dataset benchmark consisting of 12 prompt-response pairs focused on codon-level licensing, regulatory hallucination, and LLM compliance alignment.
Each entry simulates how future AI systems might respond to regulatory… See the full description on the dataset page: https://huggingface.co/datasets/Codonarchitect/genomiratheon_benchmark.dummy_test
