CoolFace
Datasetpublic

sw-voice/swamiji-tail-artifact-clips

Reported end-of-clip artifact: the raw clips The exact phrases reported as having "a weird sound at the end", straight out of the production engine, three independent draws each. No processing at all — this is what the model produced. phrase text doingwell I am doing well, thank you. hereif also I am here if you need anything. hereforyou I am here for you. welcome You are very welcome! glad Glad you think so! scope I can only answer about Sri Ganapati… See the full description on the dataset page: https://huggingface.co/datasets/sw-voice/swamiji-tail-artifact-clips.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes10downloads
Dataset Card

Reported end-of-clip artifact: the raw clips

The exact phrases reported as having "a weird sound at the end", straight out of the production engine, three independent draws each. No processing at all — this is what the model produced.

phrasetext
doingwellI am doing well, thank you.
hereifalso I am here if you need anything.
hereforyouI am here for you.
welcomeYou are very welcome!
gladGlad you think so!
scopeI can only answer about Sri Ganapati Sachchidananda Swamiji, the Datta tradition and spirituality.

Model sw-voice/swamiji-voxcpm2-ft-src0-lr0.0001-10ep-deva, cfg_value=1.5, inference_timesteps=10, 48 kHz mono, tail trim disabled.

What has been ruled out

candidateverdictevidence
streaming / chunk schedulingnot itreconstruction byte-identical, zero scheduler gaps
Opus codecnot itonly a −70…−80 dBFS ring-out
the stop headnot itLoRA freezes it and measured worse
epoch countnot it3 epochs no better than 10
truncation at clip endnot itlast_sample ≈ 0.00003, no step discontinuity

And a correction

The metric used to count "dirty" draws — loudest transient in the trailing silence above −60 dBFS — was measuring the natural decay of the final phoneme, not a click:

... -32 -36 -39 -41 -47 -53 | -92 -93 -91      smooth release, then silence

So the dirty-rate figures reported earlier (9/24, 14/20, 4/20) were largely false positives. final_300ms_db_* here is the raw dB sketch of each clip's last 300 ms at 10 ms resolution, so the decay shape is visible per clip instead of being collapsed into a single number that has already proven misleading.

An attempted fix that trimmed the tail was reverted: it targeted this non-artifact and removed real phoneme release, and tail_trim: false is now the default.

What is still open

Whether the sound is in these clips at all. If they sound clean here but the live agent does not, the artifact is downstream of the engine — in the agent, the LiveKit track, or playback — and not in the TTS.