datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
google-analogy-dataset
Google Analogy Dataset (Columnized)
This is a columnized version of the Google analogy dataset by Mikolov et al. (2013):
https://github.com/nicholas-leonard/word2vec/blob/master/questions-words.txt
The dataset contains word analogy questions grouped by subjects such as:
capital-common-countries (e.g., Athens Greece Tokyo Japan)
currency (e.g., USA dollar Japan yen)
gram3-comparative (e.g., big bigger cold colder)
The original dataset is widely used in word embedding evaluation… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/google-analogy-dataset.word2vec_analogyAdapted from https://github.com/nicholas-leonard/word2vec
story_analogyStoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding
This is the StoryAnalogy dataset in the EMNLP'23 paper: StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding.
If you use this research, please cite us:
@inproceedings{jiayang2023storyanalogy,
title={StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical… See the full description on the dataset page: https://huggingface.co/datasets/JoeyCheng/story_analogy.permuted-letter-string-analogies
Permuted Letter-String Analogies
This repository contains datasets introduced in Hellwig et al. (2026). The datasets are an extension of the letter-string analogies introduced in Lewis & Mitchell (2025).
Each folder contains a dataset with different data attributes, and contains a train, validation, and test set.
The naming convention is:
all_transformations_<copy>_study<N>_perm<N>
all_transformations:
Below are illustrations for each transformation on the standard alphabet.… See the full description on the dataset page: https://huggingface.co/datasets/philipp-hellwig/permuted-letter-string-analogies.analog-train-identifierAnalog-identifier-Train-dataset
