CoolFace
Modelpublic

ChrisColeTech/scenema-audio

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
2likes62downloads
Model Card

Scenema Audio


Sample Audio - Honey Buns (30s)

<audio controls src="https://huggingface.co/ChrisColeTech/scenema-audio/resolve/main/samples/ComfyUI_00024.mp3"></audio>

  • —~3.3 min on an RTX 5090 (32 GB)

Sample Audio - WAKE UP PEOPLE (2m)

<audio controls src="https://huggingface.co/ChrisColeTech/scenema-audio/resolve/main/samples/ComfyUI_00030.mp3"></audio>

  • —~4.3 min on an RTX 5090 (32 GB)

Components used to generate

ComponentFile~SizeDownload
Transformerscenema-audio-transformer-int8.safetensors4.91 GBLink
Text Encodergemma-3-12b-it-Q4_K_M.gguf7.3 GBLink
Audio Pipeline VAEscenema-audio-pipeline.safetensors6.71 GBLink
Audio Encoder VAEscenema-audio-vae-encoder.safetensors42.7 MBLink
Extras folderscenema-audio/extras2.3 GBLink

⚠ scenema-audio extras is required for longer audio, and for voice-to-voice ⚠

  • —Download extras folder Link and place here: /ComfyUI/models/scenema-audio/extras
  • —Create the scenema-audio/extras folder if it doesnt exist

Must have:

  • —/scenema-audio/extras/mel-band-roformer
  • —/scenema-audio/extras/bigvgan
  • —/scenema-audio/extras/campplus
  • —/scenema-audio/extras/seedvc
  • —/scenema-audio/extras/whisper-small

⚠ Must use the updated GGUF Loader ⚠

In comfy:

  • —Open the ComfyUI Manager
  • —Change the channel to "Channel (remote)"
  • —and search for comfyui-gguf-loader

image

Use the workflow Link

image


Sources