CoolFace
Datasetpublic

depinwang/jinyang-omentum-pds-nmf-subtypes-results-v1

jinyang-omentum-pds-nmf-subtypes-results-v1 NMF splicing-subtype clustering + survival analysis on the PDS-only (Treatment_strategy=='PDS') subset of the Omental-site HGSOC cohort (106 of 168 samples), reusing the exact method from /Users/depin/src/tries/2026-01-27-three_sites_analysis_from_ovarian_cancer_data_by_cursor_agent. Tests whether that project's original whole-cohort Omental finding (k=2, log-rank p=0.0002, HR=2.10, C-index=0.591, n=168 PDS+NACT mixed) is robust to… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-pds-nmf-subtypes-results-v1.

sourceHugging Facemitupdated 13d agoView on Hugging Face
0likes188downloads
Dataset Card

jinyang-omentum-pds-nmf-subtypes-results-v1

NMF splicing-subtype clustering + survival analysis on the PDS-only (Treatment_strategy=='PDS') subset of the Omental-site HGSOC cohort (106 of 168 samples), reusing the exact method from /Users/depin/src/tries/2026-01-27-three_sites_analysis_from_ovarian_cancer_data_by_cursor_agent. Tests whether that project's original whole-cohort Omental finding (k=2, log-rank p=0.0002, HR=2.10, C-index=0.591, n=168 PDS+NACT mixed) is robust to treatment-strategy heterogeneity.

Headline (see `result_summary` config for the single-row version): NUANCED POSITIVE. chosen_k=2, cluster sizes 89/17. PRIMARY omnibus log-rank (full n=106) p=0.0019, primary multivariate Cox HR=2.78 (95% CI 1.14-6.79, p=0.025, C-index=0.682) -- same direction/rough magnitude as the original HR=2.10. BUT the one-sample-per-patient SENSITIVITY check (needed: 8/96 patients contribute 2-3 samples each) is borderline non-significant: log-rank p=0.055, Cox HR=1.93 (p=0.059). No RNA-quality/sequencing-platform covariates exist for this cohort's metadata, so the platform-confound risk flagged by the already-completed, methodologically-related analysis/omentum-pds-splicing-subtypes analysis (different, re-quantified PSI matrix) cannot be directly tested here.

Do not cite this as a clean replication of the original finding without these caveats.

Robustness check (see `diagnostic_comparison_10x_maxiter`): the official canary hit 100% (50/50) NMF ConvergenceWarning at every k=2..5 at the pinned maxiterselection=200. A follow-up run with maxiterselection/final raised 10x (2000/2000, same seed/data/nruns, job 75352707) dropped non-convergence to 0/50 at every k. chosenk, cluster sizes, primary log-rank p, sensitivity log-rank p, and all 3 Cox HRs are numerically IDENTICAL between the two runs (to >=4 significant figures). Conclusion: the non-convergence caveat is resolved -- it was a max_iter artifact, not a sign of an unstable clustering solution. The nuanced-positive result stands.

Configs

ConfigRowsContents
rank_selection4Official canary (max_iter=200): per-k (2-5) NMF diagnostics -- cophenetic, silhouette, cluster sizes, non-convergence count
cluster_assignment106Official canary: per-sample subtype label, confidence, silhouette, patient_id
survival_summary5Official canary: primary + sensitivity log-rank, plus per-Cox-model n/C-index/flags
cox_models5Official canary: full lifelines .summary rows for all 3 Cox models
clinical_associations3Official canary: subtype vs. Stage/Age/Progression association tests
result_summary1Official canary: one-row headline with every key number + caveats
rank_selection_diagnostic_10x_maxiter4Diagnostic (maxiter=2000): same table as `rankselection`, re-run at 10x max_iter
survival_diagnostic_10x_maxiter5Diagnostic: same table as survival_summary, re-run at 10x max_iter
cox_models_diagnostic_10x_maxiter5Diagnostic: same table as cox_models, re-run at 10x max_iter
diagnostic_comparison_10x_maxiter1One-row official-vs-diagnostic side-by-side comparison + plain-text conclusion

Raw files (not loadable as datasets configs, for full reproducibility)

  • —raw/gate_status.json -- official canary's full nested rank-selection + final-clustering detail
  • —raw/survival_report.json -- official canary's full nested survival/Cox diagnostics

(The diagnostic run's equivalent raw JSONs are not uploaded separately -- every number in them is already captured in the *_diagnostic_10x_maxiter configs above.)

Reproducibility

json
{
  "python": "3.13.9",
  "numpy": "2.4.1",
  "pandas": "3.0.0",
  "scipy": "1.17.0",
  "scikit-learn": "1.8.0",
  "lifelines": "0.30.0",
  "base_seed": 42,
  "official_canary": {
    "n_runs_selection": 50,
    "max_iter_selection": 200,
    "n_runs_final": 200,
    "max_iter_final": 500,
    "job": "turso:75352672, short partition, cs_ukko2, 8 CPU, 16GB, completed 2026-09-15T11:38:20Z in 8m42s"
  },
  "diagnostic_10x_maxiter": {
    "n_runs_selection": 50,
    "max_iter_selection": 2000,
    "n_runs_final": 200,
    "max_iter_final": 2000,
    "job": "turso:75352707, short partition, cs_ukko2, 8 CPU, 16GB, completed 2026-09-15T12:06:56Z in 16m44s"
  },
  "tier1_min_cluster_size": 15,
  "tier2_min_cluster_size": 11,
  "k_range": [
    2,
    3,
    4,
    5
  ]
}

Pinned venv built via uv on turso (Python 3.13.9), exact version match to the sibling experiment jinyang-gse138866-nmf-subtypes's local + turso venvs.

Provenance

  • —Experiment: jinyang-omentum-pds-nmf-subtypes (notes/experiments/jinyang-omentum-pds-nmf-subtypes)
  • —Official canary job: turso:75352672, short partition, cs_ukko2, 8 CPU / 16GB, completed 2026-09-15T11:38:20Z in 8m42s, artifact status: final (canary IS the full-spec run; data is small enough that no separate scale-up is needed)
  • —Diagnostic (robustness-check) job: turso:75352707, short partition, cs_ukko2, 8 CPU / 16GB, completed 2026-09-15T12:06:56Z in 16m44s, user-approved after seeing the official canary's 100% NMF non-convergence caveat
  • —Input data: ome_psi_matrix.csv (24,319 events x 168 samples, filtered to 106 PDS samples), ome_metadata.csv -- both from /Users/depin/src/tries/2026-01-27-three_sites_analysis_from_ovarian_cancer_data_by_cursor_agent, referenced in place (not committed to git)