CoolFace
Modelpublic

cuio/CENO-1B-1m

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
1likes117downloads
Model Card

CENO-1B-1m

CENO-1B-1m is a checkpoint of the CENO base DNA foundation model (1M context (stage 4)). It is a plain causal language model over genomic sequence on a Nemotron-H Mamba/Attention/MoE hybrid backbone, with no MSA inputs.

This checkpoint is part of the CENO DNA foundation model family. The model code, VEP pipeline, and generation demo live in the sibling CENO code repository; this directory is standalone-loadable via trust_remote_code=True (the model code is bundled here).

Model details

FamilyCENO (base)
Stage1M context (stage 4)
Parameters1302.4M
Precisionbfloat16
Weightsmodel.safetensors
model_typeceno
architecturesCENOForCausalLM
auto_map → modelmodeling_ceno.CENOForCausalLM
auto_map → tokenizerceno_tokenizer.CENOCharLevelTokenizer

Architecture

| Hidden layers | 38 | | Context length | 1048576 | | Vocab size | 512 | | Attention heads | 16 | | Intermediate size | 4096 | | Num experts (MoE) | 8 | | Experts per token | 2 |

The backbone is a Mamba / Attention / Mixture-of-Experts hybrid (Nemotron-H architecture). The tokenizer is byte-level (character-level), mapping DNA characters to their ASCII byte codes (vocab size 512).

Loading

python
from transformers import AutoModelForCausalLM, AutoTokenizer

ckpt = "CENO-1B-1m"   # path to this directory
model = AutoModelForCausalLM.from_pretrained(ckpt, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(ckpt, trust_remote_code=True)

ids = tokenizer.encode("ATCGATCG", return_tensors="pt")
# out = model.generate(ids, max_new_tokens=128)   # needs a GPU (Mamba kernels)
The Mamba layers require CUDA kernels, so forward / generation needs a GPU. Config, tokenizer, and weight loading are CPU-safe.

Intended use

  • —*Base checkpoints (CENO-)**: genomic-sequence generation and embedding extraction; downstream adaptation (fine-tuning, probing) on genomics tasks.
  • —*MSA checkpoints (CENO-P-)**: variant effect prediction (VEP) by scoring wild-type vs. variant sequences with delta log-likelihood, using the MSA scoring path. See the TraitGym VEP example in the CENO code repository.

License

Apache-2.0. The bundled model code is derived from NVIDIA's Nemotron-H HuggingFace implementation (Apache-2.0); the tokenizer is derived from Arc Institute's Evo2 CharLevelTokenizer (Apache-2.0). See the LICENSE and NOTICE files in this directory for full attribution.