autoencoder
AutoencoderDataset
AutoencoderDataset
Data for training molvae — a SELFIES molecular VAE (matryoshka latent) for Bayesian
optimization of mechanophores. Two layers:
raw/ — the source chemical databases (SMILES/SELFIES, pre-tokenized), as collected.
mixes/ — the derived training datasets: weighted, shuffled, materialized token-shard
"mixes" built from the raw sources. Each mix is documented by its config + README + build
code, so it is fully reproducible. The multi-GB token shards themselves live… See the full description on the dataset page: https://huggingface.co/datasets/MechanophoresResearch/AutoencoderDataset.autoencoder-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models.
It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses).
The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.icml26-repro-nonlinear-autoencoder-pca
Nonlinear autoencoder / PCA reproduction
This bundle reproduces the paper with the authors' pinned release
(SPOC-group/advantage_nonlinearity@4378017) plus an independent direct
population-gradient-flow audit.
run_official_amp_campaign.py: 96 finite-dimensional runs of the released
two-spike AMP at d=256/512 and eight sample ratios.
run_official_ae_campaign.py: 48 full-batch Adam runs of the released tied,
one-neuron ReLU/ELU autoencoder at d=512/1024, with exact PCA and 20,000… See the full description on the dataset page: https://huggingface.co/datasets/SabaPivot/icml26-repro-nonlinear-autoencoder-pca.p2-etf-factor-autoencoder-resultslegal-ir-autoencoder-checkpoints
Legal IR Autoencoder Checkpoints
This dataset stores checkpoint artifacts for the legal text -> formal logic IR autoencoder/Codex optimization loop.
Latest checkpoint in this upload: checkpoints/20260630T221836Z/.
Contents
state/legal-ir-autoencoder-canonical.state.json: canonical feature-level autoencoder warm-start state.
reports/: weight review, deprecation manifest, and consensus feature manifests.
scripts/review_autoencoder_weight_runs.py: script used to… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/legal-ir-autoencoder-checkpoints.p2-etf-variational-autoencoder-results
