laion/moss-va-cfg-guidance
0
Classifier-free guidance on the emotion condition: for every frame the model runs twice — once with the full instruction, once with the emotion removed from it — and the difference is amplified. No retraining involved.
96 clips, 24 utterances × 4 guidance strengths (1.0 = reference, 1.5, 2.0, 3.0). Within a row the prompt and the random seed are identical, so the only difference you hear is the amplification.
Prompt text was written for this study; all audio is generated by the model.
