CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01recursal /reprocessed_singapore_national_speech_corpus Dataset Card for Reprocessed National Speech Corpus NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here. Dataset Details Dataset Description Dataset Description: The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audiotext-generation1M<n<10M7 likes362 downloads2y agoHugging Face02matsuo-lab /JP-LLM-Corpus-PII-Filtered-10B CommonCrawl Japanese (Filtered PPI) Dataset 本データセットは、CommonCrawlより抽出した約100億(10B)トークン規模の日本語テキストデータから、特に配慮が必要な「要配慮個人情報」をフィルタリング処理したものです。 データセットの概要 元データソース: CommonCrawl(https://commoncrawl.org/) トークン数: 約10Bトークン 言語: 日本語 処理内容: 要配慮個人情報をルールベースおよび機械学習分類器を用いてフィルタリング フィルタリングには以下のコードを使用しております。https://github.com/matsuolab/jp-llm-corpus-pii-filter/ 注意事項 本データセットは、非常に大規模なテキストから自動的に要配慮個人情報を除去したものであり、完全な排除を保証するものではありません。そのため、二次的な活用に際しては、目的に応じた適切な管理・配慮が必要です。… See the full description on the dataset page: https://huggingface.co/datasets/matsuo-lab/JP-LLM-Corpus-PII-Filtered-10B.text10K<n<100K1 likes143 downloads2y agoHugging Face03Aalto-Speech-Synthesis /stortinget_speech_corpus_v1.0 Dataset Card for Stortinget Speech Corpus V1.0 Overview This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability. The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.audioautomatic-speech-recognition100K<n<1M0 likes104 downloads5mo agoHugging Face04ZhuofengLi /browsecomp-plus-corpustext10K<n<100K0 likes74 downloads6mo agoHugging Face05SEIEZ /Common_Voice_Corpus_21.0text100K<n<1M0 likes7 downloads1y agoHugging Face06velkadamban /Leipzig_Corpustextn<1K0 likes6 downloads2y agoHugging Face07Confusion24 /cv-corpus-germantext100K<n<1M0 likes4 downloads11mo agoHugging Face08Confusion24 /cv-corpus-portuguesetext100K<n<1M0 likes3 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.