CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gdsu /csd_filesimagen<1K0 likes4k downloads1y agoHugging Face02ShuhongZheng /UNO1m-filtered-splitimage100K<n<1M0 likes282 downloads1y agoHugging Face03krishnakalyan3 /emo_speech_filtered_v12 second filtered emotional speech in webdataset format https://huggingface.co/datasets/EQ4You/Emotional_Speech audio10K<n<100K0 likes181 downloads2y agoHugging Face04matsuo-lab /JP-LLM-Corpus-PII-Filtered-10B CommonCrawl Japanese (Filtered PPI) Dataset 本データセットは、CommonCrawlより抽出した約100億(10B)トークン規模の日本語テキストデータから、特に配慮が必要な「要配慮個人情報」をフィルタリング処理したものです。 データセットの概要 元データソース: CommonCrawl(https://commoncrawl.org/) トークン数: 約10Bトークン 言語: 日本語 処理内容: 要配慮個人情報をルールベースおよび機械学習分類器を用いてフィルタリング フィルタリングには以下のコードを使用しております。https://github.com/matsuolab/jp-llm-corpus-pii-filter/ 注意事項 本データセットは、非常に大規模なテキストから自動的に要配慮個人情報を除去したものであり、完全な排除を保証するものではありません。そのため、二次的な活用に際しては、目的に応じた適切な管理・配慮が必要です。… See the full description on the dataset page: https://huggingface.co/datasets/matsuo-lab/JP-LLM-Corpus-PII-Filtered-10B.text10K<n<100K1 likes138 downloads1y agoHugging Face05MohammadGholizadeh /filimo-farsi-rawaudio100K<n<1M4 likes108 downloads1y agoHugging Face06filicos-data /tts_data_conversation_0711text1M<n<10M0 likes87 downloads1y agoHugging Face07filicosTHU /askaudio1M<n<10M0 likes74 downloads1y agoHugging Face08seastar105 /Emilia-YODAS-KO-filteredaudio100K<n<1M0 likes68 downloads2y agoHugging Face09filicos-data /tts_data_0720_machine3text100K<n<1M0 likes58 downloads1y agoHugging Face10Darknsu /voxceleb2-40k-part1-preprocess-all-files-separateaudio100K<n<1M0 likes58 downloads5mo agoHugging Face11filicos-data /tts_data_0720_machine4text100K<n<1M0 likes48 downloads1y agoHugging Face12filicos-data /tts_data_0720_machine5text100K<n<1M0 likes39 downloads1y agoHugging Face13filicos-data /tts_more_data_0718text10K<n<100K0 likes38 downloads1y agoHugging Face14filicos-data /tts_data_no_change_0712text100K<n<1M0 likes33 downloads1y agoHugging Face15filicos-data /tts_data_0720_machine2text100K<n<1M0 likes22 downloads1y agoHugging Face16oyyggbond /refvie_filteredimage100K<n<1M0 likes16 downloads6mo agoHugging Face17niqi /droid_filtered_20251113.tartext10K<n<100K0 likes15 downloads10mo agoHugging Face18Shailx /tar-files-19novaudio10K<n<100K0 likes15 downloads10mo agoHugging Face19Richardlsr /FileCodeBox_dbtext1K<n<10K0 likes14 downloads1y agoHugging Face20Soltani2022 /new_fileimage10K<n<100K0 likes14 downloads10mo agoHugging Face21laion /en_and_de_reference_voice_files_for_emotion_cloningaudio100K<n<1M2 likes12 downloads1y agoHugging Face22alonmiron /medicine_eng_exam_filteredtextn<1K0 likes11 downloads2y agoHugging Face23kkail8 /TAVGBench_filtered_720k_16fpstext100K<n<1M0 likes9 downloads1y agoHugging Face24ziyang06315 /laion-filteredimage100K<n<1M0 likes9 downloads8mo agoHugging Face25pcjjjj /xhs_filestext1K<n<10K0 likes8 downloads10mo agoHugging Face26tts-dataset /filtered-gol-datasetgated Filtered GOL Dataset midralab/gol-dataset をTTS(Text-to-Speech)学習用にフィルタリングしたデータセットです。 データセット概要 項目 値 総再生時間 約1,880時間 サンプル数 約120万 話者数 380人 データサイズ 約280GB 形式 WebDataset (.tar) 音声形式 FLAC (44.1kHz, モノラル) フィルタリング条件 基本フィルタ テキスト長: 3文字以上 音声長: 1秒以上、60秒未満 話者フィルタ 話者あたり5時間以上の音声データを持つ話者のみ テキストフィルタ(除外対象) 非言語テキスト(句読点のみ、空白のみなど) 顔文字 (^_^), (T_T) など 笑い表現 (笑), 文末の www 絵文字 英数字のみのテキスト 同一文字4回以上の繰り返し データ構造… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/filtered-gol-dataset.audiotext-to-speech1M<n<10M1 likes8 downloads8mo agoHugging Face27laulampaul /file_savetext10K<n<100K0 likes7 downloads1y agoHugging Face28simon123905 /Koala_36M_1_filtered_imgsimage1K<n<10K0 likes7 downloads1y agoHugging Face29laulampaul /file_save_refyoutubetext10K<n<100K0 likes6 downloads1y agoHugging Face30simon123905 /Koala_36M_1_filtered_imgs_3w-5wimage1K<n<10K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.