bghira/minimax-music3-rvq-reverse-distillation
MiniMax Music 3 RVQ Reverse-Distillation Traces Research traces generated from MiniMax Music 3. Each ZIP contains generated audio, sampled RVQ codes, teacher sampling logits, conditioning embeddings, Flow-VAE latents, and source prompt metadata. See campaign-config.json for generation settings and indexes/ for per-commit manifests. Dataset statistics As of 2026-08-16, the uploaded snapshot contains: 2,972 generated songs 91.84 hours of actual generated audio 1… See the full description on the dataset page: https://huggingface.co/datasets/bghira/minimax-music3-rvq-reverse-distillation.
MiniMax Music 3 RVQ Reverse-Distillation Traces
Research traces generated from MiniMax Music 3. Each ZIP contains generated audio, sampled RVQ codes, teacher sampling logits, conditioning embeddings, Flow-VAE latents, and source prompt metadata. See campaign-config.json for generation settings and indexes/ for per-commit manifests.
Dataset statistics
As of 2026-08-16, the uploaded snapshot contains:
- 2,972 generated songs
- 91.84 hours of actual generated audio
- 1 minute 51 seconds mean duration
- 1 minute 59 seconds median duration
- 2,365 songs (79.6%) that ended naturally with the semantic EOS code
Generation requests were capped at 180 seconds, representing 148.6 requested hours before early stopping. Actual durations are calculated from flow_latent_length * latent_hop_length / sampling_rate, using the values stored for every uploaded sample.
The canonical metadata lives in indexes/. It includes exact or legacy-default sampling parameters and deterministic train/holdout assignments. New shards sample a deterministic mix of temperature and flow-CFG settings; legacy shards are explicitly marked with the original fixed defaults.
RVQ row 0 is an un-emitted priming row; emitted semantic frame i is codes[i + 1]. The nominal Flow-VAE/semantic ratio is 3.4453125, but consumers must use the stored chunk alignment where available because chunk-boundary rounding is not uniform. Legacy indexes identify records where exact chunk spans were not captured.
For encoder training, re-encode audio.flac with the full upstream Flow-VAE. The captured flow_vae_latents are decoder-path consistency targets, not a replacement for audio-domain training inputs.
