RikkaBotan/nyan-jenny-format
Nyan Jenny Format A Japanese text-to-speech (TTS) dataset in Jenny TTS format, derived from the multilingual_toxicity_dataset (Japanese split). Toxic words identified by DeepSeek V3 Pro are replaced with 「にゃん」(nyan) before audio synthesis, producing naturally-sounding speech that masks harmful content. Dataset Summary This dataset was created through the following pipeline: Source data — The Japanese split of textdetox/multilingual_toxicity_dataset (5 k samples:… See the full description on the dataset page: https://huggingface.co/datasets/RikkaBotan/nyan-jenny-format.
1128
../
train-00000-of-00014.parquetdownload
train-00001-of-00014.parquetdownload
train-00002-of-00014.parquetdownload
train-00003-of-00014.parquetdownload
train-00004-of-00014.parquetdownload
train-00005-of-00014.parquetdownload
train-00006-of-00014.parquetdownload
train-00007-of-00014.parquetdownload
train-00008-of-00014.parquetdownload
train-00009-of-00014.parquetdownload
train-00010-of-00014.parquetdownload
train-00011-of-00014.parquetdownload
train-00012-of-00014.parquetdownload
train-00013-of-00014.parquetdownload
validation-00000-of-00001.parquetdownload
