marin-dna/rag-five-regions-v1-cds
marin-dna/rag-five-regions-v1-cds Human-anchored RAG documents for the cds region. The producing workflow owns source identities, projected coordinates, and sequence provenance at s3://oa-bolinas/snakemake/vertebrate_projection_dataset/results/rag-five-regions-v1/6b1593c274a886d20f5c0ddf3712916d446f5fed/10ba63bc375ba912909b82cda68143df4677431124fc8e15f6ac42850ce0fca6/full/rag. The publishing workflow owns public shard row mappings and release checksums at… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/rag-five-regions-v1-cds.
marin-dna/rag-five-regions-v1-cds
Human-anchored RAG documents for the cds region. The producing workflow owns source identities, projected coordinates, and sequence provenance at s3://oa-bolinas/snakemake/vertebrate_projection_dataset/results/rag-five-regions-v1/6b1593c274a886d20f5c0ddf3712916d446f5fed/10ba63bc375ba912909b82cda68143df4677431124fc8e15f6ac42850ce0fca6/full/rag. The publishing workflow owns public shard row mappings and release checksums at s3://oa-bolinas/snakemake/vertebrate_projection_dataset/results/rag-five-regions-publication-v1/54b6f936467bbc657ae88753a6f24b768bd3363d/e2af428648e45c1f7b23d938842b85325d9d864dc5335940da99d0cca3bbe0e0/full/rag/publication_provenance/cds.
Both splits contain only a string sequence column. Training has 581,256 rows; validation has 400 rows sampled from the complete chr18 training holdout. Every retained training locus contributes forward and reverse-complement rows. Each document joins available 255-bp species windows with atomic [SEQ] separators in a fixed per-row permutation. Human is included once; selected non-human representatives supply context, with at most 40 species in a document. Missing projections are omitted, genuine Ns and source letter case are retained, and reverse complementation acts within each segment. Coordinates in producer artifacts use hg38 and 0-based, half-open intervals.
The training consumer adds one BOS token and right padding to 10,240 positions, with padding targets excluded from loss. The strings here contain neither BOS nor padding. Public shard assignment is deterministic; public_row_index in the publisher's row mapping identifies each exact published document.
MarinDNA releases the processed dataset under OpenMDW 1.1. The underlying public genome assemblies and alignments retain their original source terms and attribution, recorded in the pinned source manifests. This release adds document assembly and does not relicense the underlying source assets.
