CoolFace
Datasetpublicgated

tts-dataset/filtered-gol-dataset

Filtered GOL Dataset midralab/gol-dataset をTTS(Text-to-Speech)学習用にフィルタリングしたデータセットです。 データセット概要 項目 値 総再生時間 約1,880時間 サンプル数 約120万 話者数 380人 データサイズ 約280GB 形式 WebDataset (.tar) 音声形式 FLAC (44.1kHz, モノラル) フィルタリング条件 基本フィルタ テキスト長: 3文字以上 音声長: 1秒以上、60秒未満 話者フィルタ 話者あたり5時間以上の音声データを持つ話者のみ テキストフィルタ(除外対象) 非言語テキスト(句読点のみ、空白のみなど) 顔文字 (^_^), (T_T) など 笑い表現 (笑), 文末の www 絵文字 英数字のみのテキスト 同一文字4回以上の繰り返し… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/filtered-gol-dataset.

sourceHugging Faceotherupdated 9mo agoView on Hugging Face
1likes9downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
tts-dataset/filtered-gol-dataset · CoolFace