datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UNO1m-filtered-splitemo_speech_filtered_v12 second filtered emotional speech in webdataset format
https://huggingface.co/datasets/EQ4You/Emotional_Speech
JP-LLM-Corpus-PII-Filtered-10B
CommonCrawl Japanese (Filtered PPI) Dataset
本データセットは、CommonCrawlより抽出した約100億(10B)トークン規模の日本語テキストデータから、特に配慮が必要な「要配慮個人情報」をフィルタリング処理したものです。
データセットの概要
元データソース: CommonCrawl(https://commoncrawl.org/)
トークン数: 約10Bトークン
言語: 日本語
処理内容: 要配慮個人情報をルールベースおよび機械学習分類器を用いてフィルタリング
フィルタリングには以下のコードを使用しております。https://github.com/matsuolab/jp-llm-corpus-pii-filter/
注意事項
本データセットは、非常に大規模なテキストから自動的に要配慮個人情報を除去したものであり、完全な排除を保証するものではありません。そのため、二次的な活用に際しては、目的に応じた適切な管理・配慮が必要です。… See the full description on the dataset page: https://huggingface.co/datasets/matsuo-lab/JP-LLM-Corpus-PII-Filtered-10B.Emilia-YODAS-KO-filtereddroid_filtered_20251113.tarrefvie_filteredmedicine_eng_exam_filteredlaion-filteredfiltered-gol-dataset
Filtered GOL Dataset
midralab/gol-dataset をTTS(Text-to-Speech)学習用にフィルタリングしたデータセットです。
データセット概要
項目
値
総再生時間
約1,880時間
サンプル数
約120万
話者数
380人
データサイズ
約280GB
形式
WebDataset (.tar)
音声形式
FLAC (44.1kHz, モノラル)
フィルタリング条件
基本フィルタ
テキスト長: 3文字以上
音声長: 1秒以上、60秒未満
話者フィルタ
話者あたり5時間以上の音声データを持つ話者のみ
テキストフィルタ(除外対象)
非言語テキスト(句読点のみ、空白のみなど)
顔文字 (^_^), (T_T) など
笑い表現 (笑), 文末の www
絵文字
英数字のみのテキスト
同一文字4回以上の繰り返し
データ構造… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/filtered-gol-dataset.Koala_36M_1_filtered_imgsKoala_36M_1_filtered_imgs_3w-5wTAVGBench_filtered_720k_16fpscontrolpose_filteredworldedit_text_filteredworldedit_action_filteredworldedit_color_filteredworldedit_resize_filteredworldedit_addition_filteredNoises_filtered
