marin-dna/phylop-uniform-v1-background
marin-dna/phylop-uniform-v1-background Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-background.
marin-dna/phylop-uniform-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence case is independent of that filter: lowercase bases preserve source repeat masking, uppercase bases preserve source non-repeat-masked sequence, and conservation scores never rewrite emitted characters or case.
Produced by the commit-pinned vertebrate projection pipeline.
Human anchor provenance
The authoritative projection table was produced by source commit 2162b6aa8299a9748eeb8031318b49072bb8c3fc with source config SHA-256 94d512050de327f96fda1105ce9c6ae5562944e402802516c7cde54795d8cdd1. Human anchors use the uniform GRCh38 255 bp grid at 128 bp stride and require at least 51 positions with hg38.phyloP447way >= 2.2162. The six labels are the issue #232 v4 CDS, 3-prime UTR, protein-coding TSS/5-prime UTR, and ncRNA-exon assignments; issue #326 Arm A enhancer; and the exhaustive phyloP-constrained remainder background.
Coordinates and sequence semantics
Human source coordinates use GRCh38/hg38 primary chromosomes. All human and target coordinates are 0-based and half-open. Every emitted sequence is exactly 255 bases, preserves its source FASTA/2bit letter case, and is oriented to the human anchor.
The published format is zstd-compressed JSON Lines under data/train/ and data/validation/. The release is validated before upload against a producer-keyed manifest of file sizes and SHA-256 checksums, and the immutable Hub revision exposes each large-file SHA-256 identifier.
license: other preserves the source-specific terms of the public genome assemblies, annotations, and alignments rather than relicensing those inputs.
Splits
train: 52,119,732 rows after removing the validation sample and applying the configured reverse-complement augmentation.validation: 16,384 original-orientation rows sampled uniformly without replacement with seed 517 before augmentation (4,194,304 tokens including BOS).
The split is row-level and does not stratify by chromosome, species, or human anchor. Different species projections from one human anchor may occur on opposite sides of the split. The reverse complement of a selected validation row is excluded from training.
The selected target manifest contains 135 family-deduplicated projection targets; human reference rows are added separately once per anchor.
Schema
query_name:Stringsource_chrom:Stringsource_start:Int64source_end:Int64region_label:Stringspecies:Stringalignment_name:Stringassembly:Stringtaxonomy_id:Int64family:Stringclade:Stringphylogenetic_rank:Int64alignment_source:Stringt_chrom:Stringt_start:Int64t_end:Int64t_strand:Stringt_src_size:Int64pre_resize_t_start:Int64pre_resize_t_end:Int64fragment_count:Int64aligned_bases:Int64sequence:Stringaugmentation:String
Intended use and limitations
This dataset is intended for genomic language-model research and is not a clinical resource. Assembly quality, alignment gaps, repeat masking, family-deduplicated species selection, human conservation selection, and the center-nucleotide acceptance contract affect the observed sequence distribution. Projecting the center nucleotide does not establish that both 127 bp target flanks are homologous to the full human anchor.
