CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AbstractPhil /wordnet-lexical-topology WordNet Lexical Topology Dataset Dataset Summary The WordNet Lexical Topology Dataset provides comprehensive n-gram frequency analysis from multiple sources: NLTK WordNet: Original Princeton WordNet with 117,659 synsets HF WordNet: Frequency-weighted definitions from 864,894 entries with cardinality data Unicode: Character names from 143,041 Unicode codepoints This dataset preserves sequential information crucial for language modeling and text generation, with over 12… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-lexical-topology.tabulartext-generation10M<n<100M2 likes113 downloads1y agoHugging Face02evalitahf /lexical_substitutionThe Lexical Substitution Task Test Set comprehends the test set and the gold labels used in the Lexical Substitution Task (https://www.evalita.it/2009/tasks/lexical), organised as part of the EVALITA 2009 evaluation campaign (http://www.evalita.it/2009). The task challenged participants to build systems that could automatically find synonyms for a set of 231 words appearing in different contexts. The data set contains 1710 sentences extracted from the Italian Syntactic Semantic Treebank (ISST)… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/lexical_substitution.texttext-generation1K<n<10K0 likes73 downloads2y agoHugging Face03shubhamg2208 /lexicapLexicap contains the captions for every Lex Friedman Podcast episode. It it created by [Dr. Andrej Karpathy](https://twitter.com/karpathy). There are 430 caption files available. There are 2 types of files: - large - small Each file name follows the format `episode_{episode_number}_{file_type}.vtt`.text-classificationn<1K0 likes47 downloads4y agoHugging Face04hsuvaskakoty /chew_lexical Dataset Card for Dataset Name This is the lexical/no-overlapping split of the CHEW dataset(CHEW: A Dataset of CHanging Events in Wikipedia). Dataset Details Dataset Description This dataset is the Lexical/No-overlapping split of the CHEW Dataset,where CHEW stands for CHanging Events in Wikipedia. It contains Wikipedia titles, text in two timestamped versions and Binary Label showing Change(1) or No change(0). Change here means there has been informationm… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/chew_lexical.texttext-classification1K<n<10K0 likes22 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.