datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wordnet-lexical-topology
WordNet Lexical Topology Dataset
Dataset Summary
The WordNet Lexical Topology Dataset provides comprehensive n-gram frequency analysis from multiple sources:
NLTK WordNet: Original Princeton WordNet with 117,659 synsets
HF WordNet: Frequency-weighted definitions from 864,894 entries with cardinality data
Unicode: Character names from 143,041 Unicode codepoints
This dataset preserves sequential information crucial for language modeling and text generation, with over 12… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-lexical-topology.Abstract2Appendix_v1_10k
Dataset Card: Abstract2Appendix v1
Dataset Description
The Abstract2Appendix v1 dataset is a high-quality collection of academic peer reviews and their associated research paper metadata. This dataset combines reviews from four premier machine learning and AI conferences: NeurIPS 2023, EMNLP 2023, TMLR, and ICLR 2023, shuffled into a unified corpus. It is designed to enhance long-context capabilities in Large Language Models (LLMs) and supports tasks such as fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/alexshengzhili/Abstract2Appendix_v1_10k.wordnet-definitions
WordNet Multiple Definitions - Columnar Format
Overview
This dataset is an optimized columnar version of WordNet multiple definitions, designed for high-performance queries and rapid extraction.
Each definition was sourced by GPT-5 Nano. I may update this to include additional definitions in the future, but I will not break the format.
The original dataset has a more unabridged and noisy set of data; so I'm definitely going to leave it intact. Noisy training is important… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-definitions.
