CoolFace
11 results

semcor

thesofakillers /SemCor Dataset Card for SemCor Dataset Summary SemCor 3.0 was automatically created from SemCor 1.6 by mapping WordNet 1.6 to WordNet 3.0 senses. SemCor 1.6 was created and is property of Princeton University. Some (few) word senses from WordNet 1.6 were dropped, and therefore they cannot be retrieved anymore in the 3.0 database. A sense of 0 (wnsn=0) is used to symbolize a missing sense in WordNet 3.0. The automatic mapping was performed within the Language and Information… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/SemCor.text-classification100K<n<1M2 likes484 downloads4y agoHugging FaceWSDPO /SemCortext100K<n<1M0 likes289 downloads6mo agoHugging Facespdenisov /wsd_semcor Dataset Card for "wsd_semcor" More Information needed text10K<n<100K1 likes193 downloads3y agoHugging FaceMarkChen1214 /SemCor Dataset Card for "SemCor – sense-tagged English corpus" Description This dataset is derived from the wsd_semcor dataset, originally hosted on Hugging Face. It has been preprocessed for tasks related to Word Sense Disambiguation (WSD) and WordNet integration. Preprocessing The original text data underwent the following preprocessing steps: Text splitting into individual words (lemmas). TF-IDF (Term Frequency-Inverse Document Frequency) analysis to understand… See the full description on the dataset page: https://huggingface.co/datasets/MarkChen1214/SemCor.texttext-classification10K<n<100K0 likes114 downloads3y agoHugging Facedeshanksuman /WSD_DATASET_FEWS_SEMCOR FEWS and Semcor Dataset for Word Sense Disambiguation (WSD) This repository contains a formatted and cleaned version of the FEWS and Semcor dataset, specifically arranged for model fine-tuning for Word Sense Disambiguation (WSD) tasks. Dataset Description The FEWS and Semcor dataset has been preprocessed and formatted to be directly usable for training and fine-tuning language models for word sense disambiguation. Each ambiguous word in the context is enclosed with <WSD>… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/WSD_DATASET_FEWS_SEMCOR.text100K<n<1M0 likes51 downloads2y agoHugging Facelopentu /Chinese-Wordnet-SemCor Chinese Wordnet SemCor Dataset Summary This dataset is designed for the task of Word Sense Disambiguation (WSD) for common Chinese words, specifically focusing on words identified as "difficult" (having more than 10 senses) within Chinese Wordnet (CWN) 2.0. It originates from the annotation dataset described in Section 3.1 of the paper "Resolving Regular Polysemy in Named Entities." The original dataset consisted of 28,836 example sentences where a target "difficult" word… See the full description on the dataset page: https://huggingface.co/datasets/lopentu/Chinese-Wordnet-SemCor.texttoken-classification100K<n<1M1 likes37 downloads1y agoHugging Face