CoolFace
Datasetpublic

RikkaBotan/nyan-jenny-format

Nyan Jenny Format A Japanese text-to-speech (TTS) dataset in Jenny TTS format, derived from the multilingual_toxicity_dataset (Japanese split). Toxic words identified by DeepSeek V3 Pro are replaced with 「にゃん」(nyan) before audio synthesis, producing naturally-sounding speech that masks harmful content. Dataset Summary This dataset was created through the following pipeline: Source data — The Japanese split of textdetox/multilingual_toxicity_dataset (5 k samples:… See the full description on the dataset page: https://huggingface.co/datasets/RikkaBotan/nyan-jenny-format.

sourceHugging Faceopenrail++updated 4mo agoView on Hugging Face
1likes128downloads
settings

This repository belongs to RikkaBotan on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenyan-jenny-format
visibilitypublic
licenceopenrail++
gatedno
ownerRikkaBotan
Account settings
RikkaBotan/nyan-jenny-format · CoolFace