CoolFace
Datasetpublic

depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-splice-enrichment-v1

jinyang-omentum-subtype-artifact-mechanism -- canary splicing-arm, enrichment summary One row per rMATS event type. Every value is read directly out of the job's own splicing_gates.json; nothing here is retyped by hand. Read this before quoting any number None of the enrichment results below is a scientific finding. This is a canary: chr21+chr22 only, 40 of 1048 samples. Its job was to prove the enrichment code path runs and emits a well-formed null. It did. The… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-splice-enrichment-v1.

sourceHugging Faceupdated 9d agoView on Hugging Face
0likes43downloads
Dataset Card

jinyang-omentum-subtype-artifact-mechanism -- canary splicing-arm, enrichment summary

One row per rMATS event type. Every value is read directly out of the job's own splicing_gates.json; nothing here is retyped by hand.

Read this before quoting any number

None of the enrichment results below is a scientific finding. This is a canary: chr21+chr22 only, 40 of 1048 samples. Its job was to prove the enrichment code path runs and emits a well-formed null. It did. The significant sets it produced are far too small for the restricted permutation to resolve:

AS typesignificantobserved in R-loopnull meanratiop
SE0------skipped: no significant events to test
A3SS510.3352.990.2985
A5SS200.280.001.0000
MXE2,032121111.91.080.1244
RI2665.651.060.5373

A ratio near 1.0 with p near 0.5 (MXE, RI) is what a well-behaved null looks like when it has no power. This is not evidence that batch-associated splicing events are unenriched in R-loops. At full scale the significant sets grow by roughly two orders of magnitude and the null resolves; the canary cannot speak to that.

Two defects found in the producer

1. The GC-matched shuffle crashed and no gate covers it. The secondary composition-matched analysis aborted with The truth value of an array with more than one element is ambiguous. Use a.any() or a.all() Every row carries this in gc_block_status. It is recorded in the gates JSON under gc.error and surfaced only as a WARNING in the log -- it is not a hard gate, so the job's overall: PASS does not account for it. A PASS from this job means "every hard gate passed", not "every analysis ran".

2. The `label_permutation_calibrated` gate does not test what its wording says. The gate's detail reads "median fraction significant under permuted labels must stay near 0.05", but the code enforces only a one-sided median <= 0.15 (FDR * 3), so the observed [0.0000, 0.0000, 0.0000, 0.0000, 0.0000] passes trivially. Separately, the red-team brief asks for the enrichment pipeline's distribution under permuted labels; what is implemented permutes labels through the differential test only. Both are recorded in label_perm_* columns.

The calibration is nevertheless informative in one direction: a median of 0.0 against an expected 0.05 means the Mann-Whitney p-values are conservative under the null. PSI is bounded in [0,1] and heavily tied, so the test under-detects rather than over-detects. No false-positive risk; the significant counts are lower bounds. It is also why SE returns 0 of 13,692.

The confound this contrast carries

The NovaSeq/non-NovaSeq split is entangled with treatment, not clean:

NovaSeq PDS 15 / NACT 5 non-NovaSeq NACT 13 / PDS 7

The NovaSeq arm is 75% PDS and the non-NovaSeq arm is 65% NACT. Any "batch" effect measured by this contrast is partly a treatment effect, and the full run inherits the same entanglement. This must be addressed in the analysis (stratify, or include treatment as a covariate) before any batch claim is made.

What the canary did validate

Every gate up to and including the permutation ran and passed on real data: coordinate policy pinned per AS type, chr-prefix agreement across the psi matrices / fromGTF / RLBase BED, 1048 unique matrix columns, all 40 canary samples present in the matrix with 40 distinct BAM paths, arms exactly ['NovaSeq', 'non-NovaSeq'] at the pinned 20/20, and the hand-rolled Mann-Whitney matching scipy at 0.00e+00 over 300 events for all five types.

Reproduce

scripts/canary_splice.sbatch (ePouta, arkku partition, job 2602739) -> scripts/analyze_splicing.py, --chroms chr21,chr22 --n-perm 200 --n-label-perm 20 --seed 20260918.