CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads2mo agoHugging Face02Aniket-Tathe-08 /Custom_common_voice_dataset_using_RVC Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion Custom common_voice_v11 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion) The model underwent 200 epochs of training, utilizing a total of 1 hour of audio clips. The data was scraped from Youtube. The audio in the custom generated dataset is of a YouTuber named Ajay Pandey Description license: cc0-1.0 language: - hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_common_voice_dataset_using_RVC.tabular10K<n<100K0 likes802 downloads3y agoHugging Face03KomeijiForce /CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering10K<n<100K0 likes321 downloads3y agoHugging Face04Thoria /mandarin-most-common-words-tr-en Mandarin Most Common Words (TR-EN) Overview The Mandarin Most Common Words (TR-EN) dataset is a comprehensive trilingual vocabulary resource designed for learners of Mandarin Chinese. It provides translations and practical examples in both Turkish and English, making it highly useful for bilingual education, language learning apps, and linguistic analysis. This dataset was created by Stephanie Liu and Kamil Murat Yilmaz. Dataset Content The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Thoria/mandarin-most-common-words-tr-en.text1K<n<10K2 likes150 downloads5mo agoHugging Face05Arnold /hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset tabular1K<n<10K2 likes147 downloads5y agoHugging Face06democratic-commons /gc_santegc_sante (Grande Cause Santé; Health Great Cause) is a French citizen consultation on how to act collectively for better health, prevention an well-being in France held in 2025-2026. Data This dataset contains three subsets: proposals contains the written proposals in French. Each proposal has a text content and has a unique id proposal_id. the topic and subtopic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/gc_sante.tabular100K<n<1M0 likes99 downloads16d agoHugging Face07democratic-commons /ingerenceIngerence is a French citizen consultation on how to combat information manipulation due to foreign digital interference held in 2025-2026. Data This dataset contains three subsets: proposals contains the written proposals in French. Each proposal has a text content written by an author with a unique author_id, and has a unique id proposal_id. votes contains the votes of users on propositions. Each user has a unique id user_id and votes on proposals (defined by proposal_id).… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/ingerence.tabular100K<n<1M0 likes94 downloads16d agoHugging Face08democratic-commons /eurhopeEurhope is a European-Union wide multilingual citizens consultation on the future of the European Union. It was held in 2024. Data This dataset contains two subsets: proposals contains the written proposals in French. Each proposal has a text content written by a author author_id and has a unique id proposal_id. Each proposal has one of 22 language given in its language column. votes contains the votes of users on propositions. Each user has a unique id user_id and votes on… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/eurhope.tabular100K<n<1M0 likes81 downloads16d agoHugging Face09democratic-commons /steuer_debateThe Steuer Debate consultation is a german citizen participation project on fair taxes and finances held in 2025. Data This dataset contains three subsets: proposals contains the written propositions (in german). the topic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the consultation. Each proposal has a unique id proposal_id. votes contains the votes of users on propositions. Each user has a unique id user_id and… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/steuer_debate.tabular10K<n<100K0 likes71 downloads16d agoHugging Face10mbazaNLP /common-voice-kinyarwanda-english-dataset Kinyarwanda-English Commonvoice dataset A compilation of Kinyarwanda-english dataset to be used to train multi-lingual ASR Note: The audio dataset shall be added in the future text100K<n<1M0 likes57 downloads4y agoHugging Face11dgduksict /commonvoice-mnaudio1K<n<10K0 likes47 downloads11mo agoHugging Face12bcv-commons /hebrew-lexical-references Hebrew Lexical Reference Indices Four structured, Strong's-linked transcriptions of external Hebrew (and one Hebrew↔Greek) lexical reference sources. These are not our own synonymy judgments — each config faithfully represents what an established outside source, or an actual historical translation record, already asserts (an etymological dictionary's own root groupings, a WordNet's own synset membership, five named scholars' own verified structural analysis, the Septuagint's own… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/hebrew-lexical-references.text1K<n<10K0 likes45 downloads2mo agoHugging Face13alex73 /mozilla-common-voice-23-bel-texts-exporttabulartext-generation100K<n<1M0 likes41 downloads10mo agoHugging Face14RashadGarazadeh /CommonVoiceAzaudio10K<n<100K0 likes38 downloads3y agoHugging Face15TransferRapid /CommonVoices20_ro Common Voices Corpus 20.0 (Romanian) Common Voices is an open-source dataset of speech recordings created by Mozilla to improve speech recognition technologies. It consists of crowdsourced voice samples in multiple languages, contributed by volunteers worldwide. Challenges: The raw dataset included numerous recordings with incorrect transcriptions or those requiring adjustments, such as sampling rate modifications, conversion to .wav format, and other refinements essential… See the full description on the dataset page: https://huggingface.co/datasets/TransferRapid/CommonVoices20_ro.audioautomatic-speech-recognition10K<n<100K4 likes37 downloads2y agoHugging Face16manjugeorge /commonaudio1K<n<10K0 likes36 downloads2y agoHugging Face17aranemini /commonvoicebadini Northern Kurdish (Arabic Script) ASR Dataset Dataset Description Northern Kurdish is the most widely spoken variant of the Kurdish language and is used across all parts of Kurdistan. Although it is mainly written today in the Latin script, it was historically written in the Arabic script. The Arabic script is still used for this dialect in Southern Kurdistan, particularly in the Duhok province of the Kurdistan Regional Government (KRG).Similarly, the primary writing… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/commonvoicebadini.textautomatic-speech-recognition10K<n<100K0 likes36 downloads9mo agoHugging Face18KarenSmith /common-variety-d04dad common-variety-d04dad Synthetic products test data: 45 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/KarenSmith/common-variety-d04dad.tabularn<1K0 likes35 downloads15d agoHugging Face19OpenVoiceOS /ovos-common-query-intentsDataset focusing on general knowledge questions, it distinguishs between queries aimed at a specific information source (wikipedia, wolfram alpha, duckduckgo..) and non specific queries texttext-classificationn<1K0 likes33 downloads1y agoHugging Face20alfredplpl /commoncatalog-cc-by-ext CommonCatalog CC-BY Extention このリポジトリはCommonCatalog CC-BYを拡張して、追加の情報を入れたものです。 以下の情報が追加されています。 Phi-3 VisionでDense Captioningした英語キャプション 英語キャプションをPhi-3 Mediumで日本語化した日本語キャプション 主キーはphotoidですので、CommonCatalog CC-BYと結合するなりして使ってください。 streaming=Trueで読み込むと同じ順に読み込まれますのでそれを利用するのが一番楽です。 License 画像がCC BYなため、わかりやすくCC BYにしています。したがって、商用利用可能です。 Sample Code import pandas from datasets import load_dataset df=pandas.read_csv("commoncatalog-cc-by-phi3-ja.csv") dataset =… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/commoncatalog-cc-by-ext.texttext-to-image10K<n<100K8 likes31 downloads2y agoHugging Face21AI-Culture-Commons /philosophy-culture-translations-html-csv AI-Culture Philosophy and Culture Translations CSV + HTML Corpus The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind. This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.imagetranslation1K<n<10K2 likes31 downloads1y agoHugging Face22Ember-Wisp /common-strategy-d0489c common-strategy-d0489c Synthetic sensors test data: 46 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Ember-Wisp/common-strategy-d0489c.tabularn<1K0 likes31 downloads15d agoHugging Face23CocoaRain /common_voice_13_0_zh_pseudo_labelledtext1K<n<10K0 likes23 downloads3y agoHugging Face24CS-224s /common-voicetabular10K<n<100K0 likes23 downloads2y agoHugging Face25alfredplpl /commoncatalog-cc-by-recap CommonCatalog CC-BY Recaptioning このリポジトリはCommonCatalog CC-BYを拡張して、追加の情報を入れたものです。 以下の情報が追加されています。 Phi-3 VisionでDense Captioningした英語キャプション 主キーはphotoidですので、CommonCatalog CC-BYと結合するなりして使ってください。 streaming=Trueで読み込むと同じ順に読み込まれますのでそれを利用するのが一番楽です。 Sample Code import pandas from datasets import load_dataset from tqdm import tqdm import json df=pandas.read_csv("commoncatalog-cc-by-phi3.csv") dataset = load_dataset("common-canvas/commoncatalog-cc-by",split="train",streaming=True)… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/commoncatalog-cc-by-recap.textimage-to-text100K<n<1M3 likes20 downloads2y agoHugging Face26bjak /common_voice_13_0_thai_small_pseudo_labelledtextn<1K0 likes18 downloads3y agoHugging Face27omarsou /common_voice_16_1_spanish_test_set Dataset Card for Common Voice Corpus 16 Spanish Dataset Acknowledgement The dataset belongs to COMMON VOICE MOZILLA FOUNDATION. I just uploaded the spanish test set (from HERE : https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main) Dataset Summary The Common Voice dataset consists of a unique MP3 and corresponding text file. Languages Spanish How to use The datasets library allows you to load and pre-process… See the full description on the dataset page: https://huggingface.co/datasets/omarsou/common_voice_16_1_spanish_test_set.tabular10K<n<100K1 likes17 downloads3y agoHugging Face28ycryu /dividend-common-question Dividend Common Question Dataset 📖 개요 이 데이터셋은 배당주 투자 관련 자주 묻는 질문과 답변을 Alpaca 포맷(instruction, input, output)으로 구성한 학습용 자료입니다.총 10여 개의 샘플이 포함되어 있으며, 파인튜닝 실습이나 자연어 처리 모델 학습에 활용할 수 있습니다. 📂 데이터 구조 데이터는 CSV 파일(dividend-common-question.csv)로 제공되며, 다음과 같은 열을 포함합니다: instruction: 모델에게 주는 지시문 (예: "배당성향이 높아야 좋은 건가요?") input: 지시문을 수행하는 데 필요한 추가 입력 (없으면 빈칸) output: 모델이 생성해야 하는 답변 (예: "꼭 그렇다고 할 수는 없습니다...") 예시 instruction,input,output "시가배당수익률이 높으면 좋은… See the full description on the dataset page: https://huggingface.co/datasets/ycryu/dividend-common-question.texttext-generationn<1K0 likes15 downloads6mo agoHugging Face29Lkhagvasurenam /common_voice_13_0_mn_pseudo_test_smalltextn<1K0 likes14 downloads3y agoHugging Face30alfredplpl /commoncanvas-cc-by-recap-2 CommonCatalog CC-BY Recaptioning 2 このリポジトリはCommonCatalog CC-BYを拡張して、追加の情報を入れたものです。 以下の情報が追加されています。 Florence-2-large-ftでDense Captioning (More detailed caption) した英語キャプション streaming=Trueで読み込むと同じ順に読み込まれますのでそれを利用するのが一番楽です。 Sample Code import pandas from datasets import load_dataset from tqdm import tqdm import json df=pandas.read_csv("commoncatalog-cc-by-phi3.csv") dataset = load_dataset("common-canvas/commoncatalog-cc-by",split="train",streaming=True) data_info=[] for… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/commoncanvas-cc-by-recap-2.textimage-to-text100K<n<1M0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.