marin-dna/zoonomia-v1-v4_ncrna_exon-order
bolinas-dna/zoonomia-v1-v4_ncrna_exon-order The bolinas-dna/zoonomia-v1-v4_ncrna_exon cross-mammal training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover). Same human anchors and same… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ncrna_exon-order.
bolinas-dna/zoonomia-v1-v4_ncrna_exon-order
The `bolinas-dna/zoonomia-v1-v4_ncrna_exon` cross-mammal training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors and same per-window sequences as `bolinas-dna/zoonomia-v1-v4_ncrna_exon` — only the set of target species differs. This is a species-axis subset (a row-filter on the species column), orthogonal to the region/intervals axis that defines v4_ncrna_exon. Produced by the `snakemake/zoonomia_projection_dataset` pipeline (commit `c005cbc7a236`).
Species cohort (order, 19 species)
The cohort is defined by `config/species_zoonomia_447_order_dedup.tsv` — one row per species (raw HAL leaf name in the species column). Because it is a strict subset of the 108-family v1 species set, the cross-mammal projection is reused as-is; no re-halLiftover is run.
Size
2,886,512 training samples across all JSONL.zst shards — the v4_ncrna_exon anchors projected onto the 19-species cohort, after reverse-complement augmentation. Fewer than the implicit-default 108-species `bolinas-dna/zoonomia-v1-v4_ncrna_exon` by roughly the species ratio.
Schema
Single train split of JSONL.zst shards at data/train/shard_NNNN.jsonl.zst, identical to `bolinas-dna/zoonomia-v1-v4_ncrna_exon`:
Construction
- Build the v1 cross-mammal training set and its
v4_ncrna_exonregion partition (see the pipeline README). - Filter
v4_ncrna_exonto theorderspecies cohort viamarin_dna.pipelines.projection.subset.filter_to_species(asserts the cohort is a subset of the projection's species). - RC-augment, shuffle (
seed=42), shard to JSONL, zstd-compress, upload viahf upload-large-folder.
Source code
- Pipeline: snakemake/zoonomia_projection_dataset (latest)
- Pinned to this dataset's build: commit `c005cbc7a236`
- Species cohort list: `config/species_zoonomia_447_order_dedup.tsv`
- Base region dataset: `bolinas-dna/zoonomia-v1-v4_ncrna_exon`
