marin-dna/vertebrate-v1-all
marin-dna/vertebrate-v1-all Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the all region cohort with all species scope and preserves source FASTA/2bit letter case. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence case is independent of that filter: lowercase bases preserve source repeat masking, uppercase bases preserve source non-repeat-masked… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-all.
marin-dna/vertebrate-v1-all
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the all region cohort with all species scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence case is independent of that filter: lowercase bases preserve source repeat masking, uppercase bases preserve source non-repeat-masked sequence, and conservation scores never rewrite emitted characters or case.
Produced by the commit-pinned vertebrate projection pipeline.
Splits
train: 240,128,926 rows; no chromosome-18 source anchors.validation: 16,384 original-orientation chromosome-18 rows (4,194,304 tokens including BOS).
The selected target manifest contains 135 family-deduplicated projection targets; human reference rows are added separately once per anchor.
Schema
query_name:Stringsource_chrom:Stringsource_start:Int64source_end:Int64region_label:Stringspecies:Stringalignment_name:Stringassembly:Stringtaxonomy_id:Int64family:Stringclade:Stringphylogenetic_rank:Int64alignment_source:Stringt_chrom:Stringt_start:Int64t_end:Int64t_strand:Stringt_src_size:Int64pre_resize_t_start:Int64pre_resize_t_end:Int64fragment_count:Int64aligned_bases:Int64sequence:Stringaugmentation:String
