CoolFace
Apppublic

laion/emotion-crossfade-v2-listening

sourceHugging Faceapache-2.0updated 25d agoView on Hugging Face
0likes
App README

Making a performance turn, without steering

A second round of the emotional-transition study on the LAION voice-acting model. The question is unchanged: can the model turn from one feeling into another inside two sentences, the way a person does, rather than switching between them at the sentence break?

What changed. The first round built the fade out of steering vectors — nudging the model's internal state towards one feeling and away from another. Those measure well and, a listener reported, sound wrong: more emotional, with strange artefacts, and the timbre off. The demo server has rolled its default back to adapters only. This round therefore drops steering entirely and rebuilds the fade from the three levers that are still trusted — the LoRA adapters, the brief, and classifier-free guidance — crossed against four guidance levels (off, 2, 3, 4; field use puts the usable ceiling at about 4).

Read the anchor separation first. Every number here comes from a scoring model, and the first round's central finding was about that scorer rather than about any method: the same sentence briefed as intense sadness and as intense amusement differed by 0.12 of its own scatter, and 0 of 30 texts cleared a full unit. The page opens with the same measurement made three different ways on three different segment lengths, so a reader can see for themselves how much weight the ranking underneath it will carry.

The page explains every condition in plain language and lets you rate each clip. Ratings stay in your browser; a button copies them out as text.

Base model: `laion/moss-tts-local-transformer-4.55b-voice-acting-v2-sft3`.