RikkaBotan/nyan-jenny-format
Nyan Jenny Format A Japanese text-to-speech (TTS) dataset in Jenny TTS format, derived from the multilingual_toxicity_dataset (Japanese split). Toxic words identified by DeepSeek V3 Pro are replaced with 「にゃん」(nyan) before audio synthesis, producing naturally-sounding speech that masks harmful content. Dataset Summary This dataset was created through the following pipeline: Source data — The Japanese split of textdetox/multilingual_toxicity_dataset (5 k samples:… See the full description on the dataset page: https://huggingface.co/datasets/RikkaBotan/nyan-jenny-format.
This repository belongs to RikkaBotan on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
