CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Yvthyvq /KarenSo_CantoneseRecordings_Liujgoj KarenSo_CantoneseRecordings_Liujgoj 本數據集係基於開源粵語語音庫 kakiso/KarenSo_CantoneseRecordings 進行重構與正字、拼音重映射嘅溜歌粵語(Liujgoj)語音正字 Ground Truth 數據集。主要用於音標切分、粵語語音流形(Language Manifold)對齊、以及 Stage 2 SFT 翻譯與語言工程任務。 👥 致謝與上游數據集說明 (Acknowledgment & Upstream Source) 本數據集嘅原始音頻與文本來源於 Hugging Face 社群成員 kakiso 分享嘅項目: 原始數據集 (Original Dataset): kakiso/KarenSo_CantoneseRecordings 原始授權協議 (License): CC-BY-4.0 在此由衷感謝原創作者 Karen So 及其團隊錄製並無私分享高品質(44.1kHz / 16bit / Mono)嘅純淨粵語口語語料,為廣東話開源 AI… See the full description on the dataset page: https://huggingface.co/datasets/Yvthyvq/KarenSo_CantoneseRecordings_Liujgoj.audioautomatic-speech-recognition0 likes162 downloads3mo agoHugging Face02karenholzkopf /Forex_Factory_Calendar 📅 Forex Factory Economic Calendar Dataset (2007-01-01 to 2025-04-07) This dataset contains a comprehensive archive of macroeconomic calendar events sourced from Forex Factory, spanning from January 1, 2007 to April 7, 2025.Each row captures a specific event with detailed metadata including currency, event type, market impact level, reported values, and descriptive context. 📦 Dataset Summary Total timespan: 2007-01-01 → 2025-04-07 Format: CSV (UTF-8)… See the full description on the dataset page: https://huggingface.co/datasets/karenholzkopf/Forex_Factory_Calendar.texttime-series-forecasting10K<n<100K0 likes149 downloads4mo agoHugging Face03karenwky /pet-health-symptoms-dataset Pet Health Symptoms Dataset Overview This dataset contains 2,000 LLM-generated pet health symptoms text samples covering 5 common pet health condition categories, designed to train ML models for automated pet health classification. Each entry is labeled with: Pet health condition (1 of 5 distinct classes) Record type (Owner Observation or Clinical Notes) Owner observations are expressed in everyday language (e.g., "My cat scratches constantly"), whereas clinical… See the full description on the dataset page: https://huggingface.co/datasets/karenwky/pet-health-symptoms-dataset.texttext-classification1K<n<10K7 likes123 downloads1y agoHugging Face04zuleo /karen-fukuhara Karen Fukuhara textual inversion This is an embedding of Karen Fukuhara. She plays several different amazing roles from acting to voice acting: (Kimiko Miyashiro, Glimmah, Kipo). Embedding Usage Use the token kfukvf-1990 🎶 Prompt Examples 🧾 Perfectly-centered portrait-photograph of kfukvf-1990, dressed as a queen with glimmering jewelry, lifelike, subsurface scattering, photorealism, 8k resolution, beautiful, dynamic lighting ⛔ Negative prompt:… See the full description on the dataset page: https://huggingface.co/datasets/zuleo/karen-fukuhara.imagen<1K0 likes120 downloads4y agoHugging Face05Karen-Williams /spiritual-development-4b59d7 spiritual-development-4b59d7 Synthetic products test data: 34 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/Karen-Williams/spiritual-development-4b59d7.tabularn<1K0 likes41 downloads15d agoHugging Face06freococo /sagaw_karen_asrThis is the first public Sagaw Karen language ASR dataset in AI history. Sagaw Karen ASR This dataset contains audio recordings and aligned metadata in the Sagaw Karen language (ISO 639-3: ksw), a major Sgaw Karenic language spoken throughout southern and eastern Myanmar. The language is sometimes also referred to as Sgaw Karen or Sakaw Karen in English transliterations. All audio segments in this dataset were sourced from publicly available news broadcasts published by PVTV… See the full description on the dataset page: https://huggingface.co/datasets/freococo/sagaw_karen_asr.audioautomatic-speech-recognition1K<n<10K0 likes39 downloads1y agoHugging Face07freococo /karenni_language_asr_audio RFA Karenni (Kayah) Language Voices This dataset contains 17 hours of audio in the Karenni (Kayah) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Karenni language family, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. This dataset was created by freococo. The audio has been automatically segmented… See the full description on the dataset page: https://huggingface.co/datasets/freococo/karenni_language_asr_audio.audioautomatic-speech-recognition1K<n<10K0 likes37 downloads1y agoHugging Face08KarenSmith /common-variety-d04dad common-variety-d04dad Synthetic products test data: 45 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/KarenSmith/common-variety-d04dad.tabularn<1K0 likes35 downloads15d agoHugging Face09Azure-Karen /grand-mouth-c49a18 grand-mouth-c49a18 Synthetic sensors test data: 45 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Azure-Karen/grand-mouth-c49a18.tabularn<1K0 likes35 downloads15d agoHugging Face10eduhk-compling /KarenSo Dataset Description This dataset includes 50 sentences spoken in colloquial Hong Kong Cantonese (HKC), covering interrogatives and statements. Sentences are sourced from online Cantonese teaching materials and classical commercial slogans. It includes approxiametly 198 seconds of audio recorded by a female native speaker of HKC. The sampling rate is 44.1 kHz with 16-bit resolution. Transcription in Jyutping were also provided. Issues Encountered & Solution I… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/KarenSo.audion<1K0 likes34 downloads8mo agoHugging Face11CyberHarem /houjou_karen_idolmastercinderellagirls Dataset of houjou_karen/北条加蓮 (THE iDOLM@STER: Cinderella Girls) This is the dataset of houjou_karen/北条加蓮 (THE iDOLM@STER: Cinderella Girls), containing 500 images and their tags. The core tags of this character are brown_hair, brown_eyes, long_hair, breasts, bangs, medium_breasts, which are pruned in this dataset. Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization). List of… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/houjou_karen_idolmastercinderellagirls.text-to-image1K<n<10K0 likes31 downloads3y agoHugging Face12freococo /western_poe_karen_asrThis is the first public Western Poe Karen language ASR dataset in AI history. Western Poe Karen ASR This dataset contains audio recordings and aligned transcriptions in the Western Poe Karen language (also known in linguistic literature as Western Pwo or Delta Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in the Ayeyarwady Delta region of Myanmar. Although linguists commonly refer to this language as Western Pwo Karen, the community and this project prefer the spelling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/western_poe_karen_asr.audioautomatic-speech-recognition1K<n<10K0 likes30 downloads1y agoHugging Face13karen-thompson /strong-blue-dcb3f9 strong-blue-dcb3f9 Synthetic weather test data: 42 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/karen-thompson/strong-blue-dcb3f9.tabularn<1K0 likes28 downloads15d agoHugging Face14grimulkan /bluemoon_Karen_cleanedThis is a version of bluemoon-fandom-1-1-rp-cleaned further cleaned up using Karen_TheEditor 13B, in Fastchat format. I tried to fix as many of the grammatical issues as possible and didn't drop any conversations, but there are still issues since Karen is not perfect (and only 13B). If I detected any large deviations from the original text in the corrections, I fell back to a standard spell-checker, excluding estimated proper nouns from the spell-checker (which is also not perfect). I intended… See the full description on the dataset page: https://huggingface.co/datasets/grimulkan/bluemoon_Karen_cleaned.text1K<n<10K8 likes26 downloads3y agoHugging Face15freococo /eastern_poe_karen_asrThis is the first public Eastern Poe Karen language ASR dataset in AI history. Eastern Poe Karen ASR This dataset contains audio recordings and aligned metadata in the Eastern Poe Karen language (a regional variety of Eastern Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in Mon State and Kayin State in southeastern Myanmar. While linguistically described as Eastern Pwo Karen, the community and this project prefer the term Poe as a community-endorsed spelling. All audio… See the full description on the dataset page: https://huggingface.co/datasets/freococo/eastern_poe_karen_asr.audioautomatic-speech-recognition1K<n<10K0 likes24 downloads1y agoHugging Face16CyberHarem /shinomiya_karen_theidolmstermillionlive Dataset of shinomiya_karen/篠宮可憐/시노미야카렌 (THE iDOLM@STER: Million Live!) This is the dataset of shinomiya_karen/篠宮可憐/시노미야카렌 (THE iDOLM@STER: Million Live!), containing 71 images and their tags. The core tags of this character are long_hair, blonde_hair, blue_eyes, breasts, large_breasts, which are pruned in this dataset. Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/shinomiya_karen_theidolmstermillionlive.text-to-imagen<1K0 likes23 downloads3y agoHugging Face17karenlu653 /dialect_model_demotabularaudio-classificationn<1K0 likes20 downloads1y agoHugging Face18collectivat /ladino-karen-TTS Ladino Text-to-Speech (TTS) Training Dataset Dataset Description This dataset contains a single-speaker speech corpus in Ladino (Judeo-Spanish) recorded by a native speaker from Istanbul. The corpus was created for training text-to-speech synthesis models for this endangered language. Dataset Statistics Speaker: Karen (native Ladino speaker) Recordings: 1987 segments Total Duration: ~3.3 hours Sampling Rate: 16 kHz Audio Format: WAV (16-bit, mono) Language:… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/ladino-karen-TTS.audiotext-to-speech1K<n<10K0 likes20 downloads11mo agoHugging Face19PJMixers /grimulkan_bluemoon_Karen_cleaned-carded-formattedJust a simple text replace of the tags. First Character: The Beast Second Character: Belle First Character Description: A mysterious and intimidating figure, resembling a beast with a cape swishing behind him. He has an imposing presence, which he uses to assert dominance over others in his castle. His personality is stern and authoritative; he is not afraid to enforce rules or punish those who disobey him. Despite this harsh exterior, The Beast also displays signs of vulnerability and… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers/grimulkan_bluemoon_Karen_cleaned-carded-formatted.text1K<n<10K0 likes19 downloads3y agoHugging Face20karenliu /ADL_2023_HW1texttext-classification10K<n<100K0 likes17 downloads3y agoHugging Face21kakiso /KarenSo_CantoneseRecordings Dataset Description This dataset includes 50 sentences spoken in colloquial Hong Kong Cantonese (HKC), covering interrogatives and statements. Sentences are sourced from online Cantonese teaching materials and classical commercial slogans. It includes approxiametly 198 seconds of audio recorded by a female native speaker of HKC. The sampling rate is 44.1 kHz with 16-bit resolution. Transcription in Jyutping were also provided. Issues Encountered & Solution I… See the full description on the dataset page: https://huggingface.co/datasets/kakiso/KarenSo_CantoneseRecordings.audion<1K0 likes17 downloads5mo agoHugging Face22KarenVin /Footsteps_SoundEffect0 likes16 downloads2y agoHugging Face23open-llm-leaderboard-old /details_FPHam__Karen_TheEditor_V2_STRICT_Mistral_7B Dataset Card for Evaluation run of FPHam/Karen_TheEditor_V2_STRICT_Mistral_7B Dataset Summary Dataset automatically created during the evaluation run of model FPHam/Karen_TheEditor_V2_STRICT_Mistral_7B on the Open LLM Leaderboard. The dataset is composed of 1 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_FPHam__Karen_TheEditor_V2_STRICT_Mistral_7B.0 likes15 downloads3y agoHugging Face24CyberHarem /karen_granbluefantasy Dataset of karen/カレン (Granblue Fantasy) This is the dataset of karen/カレン (Granblue Fantasy), containing 24 images and their tags. The core tags of this character are hair_ornament, long_hair, brown_hair, blue_eyes, breasts, braid, large_breasts, which are pruned in this dataset. Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization). List of Packages Name Images Size… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/karen_granbluefantasy.text-to-imagen<1K0 likes13 downloads3y agoHugging Face25aravdash /karenTTSaudio1K<n<10K1 likes13 downloads1y agoHugging Face26Gopher-Lab /huberman_lab_Dr._Karen_Parker_The_Causes__Treatments_for_Autismtextn<1K1 likes10 downloads2y agoHugging Face27karenli1 /nlp-final-project-activations-3steptabular1K<n<10K0 likes9 downloads5mo agoHugging Face28Gopher-Lab /huberman_lab_Dr__Karen_Parker_The_Causes__Treatments_for_Autismtextn<1K0 likes8 downloads2y agoHugging Face29karenwhiteg /Python_Like_A_Protext10K<n<100K1 likes8 downloads2y agoHugging Face30karenlu653 /dialect_model_data shanghai-binary dataset Train/test splits for Shanghai vs Not-Shanghai binary classification. Contents data/train.parquet data/test.parquet Each row contains: audio: float array (mono) sampling_rate: 16000 dialect/label: label (Shanghai=1, else 0) tabular1K<n<10K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.