duke-trust-lab/human-aligned-similarity-benchmark
Human Aligned Similarity Benchmark You are welcome to go to alignedmachine.com to contribute. Overview This dataset contains human-aligned similarity judgments for embedding text and multimodal AI model evaluation. The benchmark is designed to assess how well AI models align with human cognitive preferences in similarity perception across text and image modalities. Dataset Structure Concept Files This dataset contains human preference… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/human-aligned-similarity-benchmark.
Human Aligned Similarity Benchmark
You are welcome to go to alignedmachine.com to contribute.
Overview
This dataset contains human-aligned similarity judgments for embedding text and multimodal AI model evaluation. The benchmark is designed to assess how well AI models align with human cognitive preferences in similarity perception across text and image modalities.
Dataset Structure
Concept Files
This dataset contains human preference judgments for word pair comparisons across three modalities:
- `concepts_tt.jsonl`: Human judgments on word pair similarities
- `concepts_it.jsonl`: Human preferences for image-text associations
- `concepts_ii.jsonl`: Human similarity judgments for image pairs
Test Data
- `v1-test.jsonl`: Human preference judgments from 115 validated users
Modality Notation
<img>prefix indicates an image concept- Text concepts have no prefix
- Modality types:
tt: Both concepts are textit: One concept is image, one is textii: Both concepts are images
Citation
If you use this dataset, please cite:
@dataset{human_aligned_similarity_benchmark_2025},
title={Human Aligned Similarity Benchmark},
author={Jiechen Li, Hannah Groos, Brinnae Bent},
year={2025},
url={https://huggingface.co/datasets/duke-trust-lab/human-aligned-similarity-benchmark}
}