words
Datasets
All datasets matching “words”polysemous-words
Polysemous Words
Polysemous Words is a large-scale collection of contextual examples for 200 common polysemous English words. Each word in this dataset has multiple, distinct senses and appears across a wide variety of natural web text contexts, making the dataset ideal for research in word sense induction (WSI), word sense disambiguation (WSD), and probing the contextual understanding capabilities of large language models (LLMs).
Dataset Overview
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/tsivakar/polysemous-words.ml_spoken_wordsMultilingual Spoken Words Corpus is a large and growing audio dataset of spoken
words in 50 languages collectively spoken by over 5 billion people, for academic
research and commercial applications in keyword spotting and spoken term search,
licensed under CC-BY 4.0. The dataset contains more than 340,000 keywords,
totaling 23.4 million 1-second spoken examples (over 6,000 hours). The dataset
has many use cases, ranging from voice-enabled consumer devices to call center
automation. This dataset is generated by applying forced alignment on crowd-sourced sentence-level
audio to produce per-word timing estimates for extraction.
All alignments are included in the dataset.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.SPEEED_s3_words_french_100k-300kSPEEED_s3_words_russian_0k-300kSPEEED_s3_words_spanish_1200k-1400k
Multilingual Audio Alignments - Processed (Mixed Text/Phonemes)
This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (spanish).
Curriculum Learning
This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule:
p_start: 0.0 (starting probability of using phonemes)
p_end: 0.0 (ending probability of using phonemes)
curriculum_rows: 400000 (rows over which probability increases)
Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_spanish_1200k-1400k.
