RareConcepts/soad-mm3-vanilla-20260822-10k-adamw8e-5-randwin128-nocapdrop
RareConcepts/soad-mm3-vanilla-20260822-10k-adamw8e-5-randwin128-nocapdrop
This is a PEFT LoRA derived from MiniMaxAI/MiniMax-Music3.
Validation was disabled during training.
The text encoder was not trained. You may reuse the base model text encoder for inference.
Training settings
- Training epochs: 416
- Training steps: 10000
- Learning rate: 8e-05
- Learning rate schedule: cosine
- Warmup steps: 50
- Max grad value: 1.0
- Effective batch size: 1
- Micro-batch size: 1
- Gradient accumulation steps: 1
- Number of GPUs: 1
- Gradient checkpointing: True
- Prediction type: autoregressivenexttoken
- Optimizer: adamw_bf16
- Trainable parameter precision: Pure BF16
- Base model precision:
no_change - Caption dropout probability: 0.0%
- LoRA Rank: 64
- LoRA Alpha: None
- LoRA Dropout: 0.1
- LoRA initialisation style: default
- LoRA mode: Standard
Training modes
- MiniMax Music train component:
language_model (global LM / RVQ planner) - MiniMax Music LM max frames:
128 - MiniMax Music LM window mode:
random
Datasets
soad-mm3-vanilla-20260822-10k-adamw8e-5-randwin128-nocapdrop-audio
- Repeats: 0
- Total number of audio files: 24
- Duration buckets: audio (24)
- Used for regularisation data: No
Inference
import torch
from diffusers import DiffusionPipeline
model_id = 'MiniMaxAI/MiniMax-Music3'
adapter_id = 'RareConcepts/soad-mm3-vanilla-20260822-10k-adamw8e-5-randwin128-nocapdrop'
pipeline = DiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.bfloat16) # loading directly in bf16
pipeline.load_lora_weights(adapter_id)
prompt = "An astronaut is riding a horse through the jungles of Thailand."
negative_prompt = 'blurry, cropped, ugly'
## Optional: quantise the model to save on vram.
## Note: The model was not quantised during training, so it is not necessary to quantise it during inference time.
#from optimum.quanto import quantize, freeze, qint8
#quantize(pipeline.transformer, weights=qint8)
#freeze(pipeline.transformer)
pipeline.to('cuda' if torch.cuda.is_available() else 'mps' if torch.backends.mps.is_available() else 'cpu') # the pipeline is already in its target precision level
model_output = pipeline(
prompt=prompt,
negative_prompt=negative_prompt,
num_inference_steps=30,
generator=torch.Generator(device='cuda' if torch.cuda.is_available() else 'mps' if torch.backends.mps.is_available() else 'cpu').manual_seed(42),
width=256,
height=256,
guidance_scale=7.5,
).images[0]
model_output.save("output.png", format="PNG")
<!-- BEGIN SOAD MM3 TOURNAMENT -->
Controlled checkpoint and strength tournament
This gallery compares Vanilla, AdamW 8e-5, random window 128, no caption dropout checkpoints under four caption conditions. Every render uses the same unseen lyrics and random stream, 40 seconds maximum duration, 30 flow steps, flow CFG 1.7, and seed 20260822. Within each batch, only LoRA strength changes.
The four caption controls separate trigger response from descriptive style learning and unrelated-style leakage:
- Off-genre, no trigger:
Alternative pop-jazz, 102 BPM, brushed drums, upright bass, clean electric piano, airy male vocals, smooth verse-chorus songwriting, restrained dynamics, polished intimate mix. - Trigger only:
system of a down, soad style music - Descriptors, no trigger:
Lilting hypnotic arpeggios, a gentle melodic mid-tempo groove with layered harmonies, sudden explosive distorted riffs, unstable meter changes, theatrical male vocals, tightly syncopated bass and drums, abrupt shifts between fragile verses and frantic heavy choruses. - Trigger + descriptors:
system of a down, soad style music, lilting hypnotic arpeggios, a gentle melodic mid-tempo groove with layered harmonies, sudden explosive distorted riffs, unstable meter changes, theatrical male vocals, tightly syncopated bass and drums, abrupt shifts between fragile verses and frantic heavy choruses.
Machine-readable manifest
checkpoint-1000
checkpoint-2500
checkpoint-4000
checkpoint-5000
checkpoint-6000
checkpoint-7000
checkpoint-8000
checkpoint-9000
checkpoint-10000
<details> <summary>Unseen lyrics used for every render</summary>
[Intro]
Count the teeth in the traffic light
Paint the silence ultraviolet
[Verse 1]
Paper crowns on the factory floor
Everybody wants a little more
Wind-up prophets in a checkout line
Selling borrowed seconds back as time
[Pre-Chorus]
Turn the dial, divide the sky
Ask the wires to tell us why
[Chorus]
Wake up the walls, let the ceiling shake
We built a feast from a plastic lake
Hands in the air while the engines stall
Who signed the order to swallow it all?
[Verse 2]
Velvet sirens in the avenue
Paint every warning a pleasant blue
I fed my shadow to the voting machine
It came back polished, quiet, and clean
[Bridge]
Softly, softly, count to four
Kick down the arithmetic door
One for the hunger, two for the show
Three for the ones who already know
[Final Chorus]
Wake up the walls, let the ceiling shake
We built a feast from a plastic lake
Hands in the air while the engines stall
Who signed the order to swallow it all?
[Outro]
Count the teeth
Kill the light
Carry the silence
Into the night</details>
NextLat predictor and XM route-embedding sidecars are training-only and are intentionally not used during inference. <!-- END SOAD MM3 TOURNAMENT -->
