CoolFace
Datasetpublic

Kaileh57/opacity-marginalization

Amortized opacity marginalization: data Training sets, trained models, priors and held-out evaluation sets for Amortized Opacity Marginalization Improves C/O Interval Calibration for Brown-Dwarf Retrievals (Heraty 2026, arXiv:2609.01665). Code and paper source: github.com/kaileh57/opacity-marginalization. Contents spectra/v7_opmarg: one million simulated NIRSpec G395H spectra generated with randomized molecular opacities (the marginalized training set)… See the full description on the dataset page: https://huggingface.co/datasets/Kaileh57/opacity-marginalization.

sourceHugging Facecc-by-4.0updated 22d agoView on Hugging Face
0likes3kdownloads
Dataset Card

Amortized opacity marginalization: data

Training sets, trained models, priors and held-out evaluation sets for Amortized Opacity Marginalization Improves C/O Interval Calibration for Brown-Dwarf Retrievals (Heraty 2026, arXiv:2609.01665). Code and paper source: github.com/kaileh57/opacity-marginalization.

Contents

  • —spectra/v7_opmarg: one million simulated NIRSpec G395H spectra generated with randomized molecular opacities (the marginalized training set)
  • —spectra/v6_cloud: one million matched spectra at fixed nominal opacities (the control training set)
  • —spectra/ablation_train: the one-axis training sets (discrete_only, smooth_only) behind the training-side ablation
  • —spectra/v7_opmarg_validation, spectra/v6_cloud_val: validation splits
  • —spectra/ablation_test, spectra/broadening_test, spectra/opmarg_ampsweep_v2: the one-axis, broadening-stress and amplitude-robustness evaluation sets
  • —heldout/clean, heldout/perturbed: the held-out sets behind every reported coverage number
  • —models/: trained normalizing-flow posterior estimators, comprising three marginalized seeds, three control seeds, three mask-augmented seeds, and the ablation and leave-one-out variants
  • —real_inputs/: preprocessed JWST NIRSpec G395H arrays (observed flux, error, bad-pixel mask per object) used for the real-data results
  • —results_gcs/: evaluation outputs from the final, amplitude-robustness, broadening and recalibration-benchmark runs
  • —priors/interlist_ratio_priors_v1.json: the opacity-perturbation prior

Spectra are HDF5 shards of 128 spectra each, with parameters, flux and noise-free flux per shard. Models are PyTorch checkpoints; loading them requires zuko (see the code repository).

License

CC BY 4.0.

Citation

See the code repository for the BibTeX entry.