CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WSDPO /SemCortext100K<n<1M0 likes252 downloads6mo agoHugging Face02spdenisov /wsd_semcor Dataset Card for "wsd_semcor" More Information needed text10K<n<100K1 likes186 downloads3y agoHugging Face03MarkChen1214 /SemCor Dataset Card for "SemCor – sense-tagged English corpus" Description This dataset is derived from the wsd_semcor dataset, originally hosted on Hugging Face. It has been preprocessed for tasks related to Word Sense Disambiguation (WSD) and WordNet integration. Preprocessing The original text data underwent the following preprocessing steps: Text splitting into individual words (lemmas). TF-IDF (Term Frequency-Inverse Document Frequency) analysis to understand… See the full description on the dataset page: https://huggingface.co/datasets/MarkChen1214/SemCor.texttext-classification10K<n<100K0 likes98 downloads3y agoHugging Face04deshanksuman /WSD_DATASET_FEWS_SEMCOR FEWS and Semcor Dataset for Word Sense Disambiguation (WSD) This repository contains a formatted and cleaned version of the FEWS and Semcor dataset, specifically arranged for model fine-tuning for Word Sense Disambiguation (WSD) tasks. Dataset Description The FEWS and Semcor dataset has been preprocessed and formatted to be directly usable for training and fine-tuning language models for word sense disambiguation. Each ambiguous word in the context is enclosed with <WSD>… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/WSD_DATASET_FEWS_SEMCOR.text100K<n<1M0 likes48 downloads2y agoHugging Face05lopentu /Chinese-Wordnet-SemCor Chinese Wordnet SemCor Dataset Summary This dataset is designed for the task of Word Sense Disambiguation (WSD) for common Chinese words, specifically focusing on words identified as "difficult" (having more than 10 senses) within Chinese Wordnet (CWN) 2.0. It originates from the annotation dataset described in Section 3.1 of the paper "Resolving Regular Polysemy in Named Entities." The original dataset consisted of 28,836 example sentences where a target "difficult" word… See the full description on the dataset page: https://huggingface.co/datasets/lopentu/Chinese-Wordnet-SemCor.texttoken-classification100K<n<1M1 likes39 downloads1y agoHugging Face06gguichard /wsd_fr_wngt_semcor_translated_aligned_all_v1 Dataset Card for "wsd_fr_wngt_semcor_translated_aligned_all_v1" More Information needed text100K<n<1M0 likes32 downloads3y agoHugging Face07gguichard /wsd_fr_wngt_semcor_translated_aligned_v2 Dataset Card for "wsd_fr_wngt_semcor_translated_aligned_v2" More Information needed text100K<n<1M0 likes25 downloads3y agoHugging Face08gguichard /wsd_fr_wngt_semcor_translated_aligned Dataset Card for "wsd_fr_wngt_semcor_translated_aligned" More Information needed text100K<n<1M0 likes12 downloads3y agoHugging Face09lavallone /selection_semcortext100K<n<1M0 likes11 downloads2y agoHugging Face10lavallone /generation_semcortext100K<n<1M0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.