CoolFace
20 results

words

tsivakar /polysemous-words Polysemous Words Polysemous Words is a large-scale collection of contextual examples for 200 common polysemous English words. Each word in this dataset has multiple, distinct senses and appears across a wide variety of natural web text contexts, making the dataset ideal for research in word sense induction (WSI), word sense disambiguation (WSD), and probing the contextual understanding capabilities of large language models (LLMs). Dataset Overview This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/tsivakar/polysemous-words.text100M<n<1B1 likes31k downloads1y agoHugging FaceMLCommons /ml_spoken_wordsMultilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.0. The dataset contains more than 340,000 keywords, totaling 23.4 million 1-second spoken examples (over 6,000 hours). The dataset has many use cases, ranging from voice-enabled consumer devices to call center automation. This dataset is generated by applying forced alignment on crowd-sourced sentence-level audio to produce per-word timing estimates for extraction. All alignments are included in the dataset.audio-classification10M<n<100M37 likes2.5k downloads4y agoHugging FaceClusterlabAi /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.texttext-generation10M<n<100M73 likes1.9k downloads2y agoHugging FaceAdoCleanCode /SPEEED_s3_words_french_100k-300ktext100K<n<1M0 likes941 downloads8mo agoHugging FaceAdoCleanCode /SPEEED_s3_words_russian_0k-300ktext100K<n<1M0 likes893 downloads7mo agoHugging FaceAdoCleanCode /SPEEED_s3_words_spanish_1200k-1400k Multilingual Audio Alignments - Processed (Mixed Text/Phonemes) This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (spanish). Curriculum Learning This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule: p_start: 0.0 (starting probability of using phonemes) p_end: 0.0 (ending probability of using phonemes) curriculum_rows: 400000 (rows over which probability increases) Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_spanish_1200k-1400k.text100K<n<1M0 likes805 downloads7mo agoHugging Face