datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
daily-papers-embeddingsCore-AlphaEarth-Embeddings
Major TOM Core AlphaEarth Embeddings Subset
This is a prototype dataset. It only includes some of the AlphaEarth embeddings stored in Major TOM grid cells.
This dataset is mostly aimed at experimentation and prototyping. It is particularly useful to use it along other datasets published within the Major TOM project.
Content
Field
Type
Description
grid_cell
string
Major TOM cell
year
int
year of the sample
thumbnail
image
3-dimensional PCA… See the full description on the dataset page: https://huggingface.co/datasets/Major-TOM/Core-AlphaEarth-Embeddings.multilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.Caselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.paired-llama-3.2-1b-embeddings-lmsys-chat-1m
Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M)
This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations.
This dataset was built to study things like:
Learning different basis for activations at a given layer
Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.embeddingsCaselaw_Access_Project_embeddingsOriginal Repository:
https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/
This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.arxiv-titles-instructorxl-embeddings
arxiv-titles-instructorxl-embeddings
This dataset contains 768-dimensional embeddings generated from the arxiv
paper titles using InstructorXL model. Each
vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The
dataset was created using precomputed embeddings exposed by the Alexandria Index.
Generation process
The embeddings have been generated using the following instruction:
Represent the Research Paper title for retrieval;… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings.platonic-embeddingsopenalex-multilingual-embeddings
OpenAlex Multilingual Embeddings
This dataset contains multilingual text embeddings of all records in OpenAlex with a title or an abstract from the snapshot of 2023-10-20.
The dataset was created for the FORAS project to investigate the efficacy of
different methods of searching in databases of academic publications. All scripts will be available in a GitHub repository.
The project is supported by a grant from the Dutch Research Council (grant no. 406.22.GO.048)… See the full description on the dataset page: https://huggingface.co/datasets/GlobalCampus/openalex-multilingual-embeddings.vision-adapter-embeddings
Vision Adapter MoonViT Embeddings
Precomputed visual embeddings used to train lightweight vision→LLM projectors
without re-running a vision tower: each row is the frozen MoonViT-V2 output
for one training image, stored as raw bfloat16 bytes.
Shards: 103 Parquet files (data/emb_0000.parquet … data/emb_0102.parquet),
1360 rows each, ~139k rows total, ~1.9 TB.
Schema per row:
column
type
meaning
key
string
embedding id, embeddings/<sha1[:20]>.pt; matches emb in… See the full description on the dataset page: https://huggingface.co/datasets/keypa/vision-adapter-embeddings.embeddings-fine-tuning
Overview
This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version.
This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning.benchmark-embeddings
Marqo Benchmark Embeddings
This dataset contains a large collection of embeddings from popular models on benchmark datasets. In addition to this, the passage data also includes the local intrinsic dimensionality (LID) for every vector considering its exact nearest 100 neighbours, LID is calculated using a Maximum Likelihood Estimation based approach.
Below is a list of the datasets and the models, every datasets queries and passages are embedded with every model.… See the full description on the dataset page: https://huggingface.co/datasets/Marqo/benchmark-embeddings.PATHOS-PLM-EMBEDDINGS
PATHOS PLM Embeddings
Precomputed protein language model (PLM) embeddings for missense substitutions and wild-type residues in 20,416 human SwissProt proteins. These embeddings are used by PATHOS to predict the pathogenicity of missense mutations.
Paper: http://dx.doi.org/10.1016/j.ailsci.2026.100165
Dataset Structure
The repository contains two config families for each PLM:
Mutation configs: <model> stores embeddings for generated missense substitutions.
Wild-type… See the full description on the dataset page: https://huggingface.co/datasets/DSIMB/PATHOS-PLM-EMBEDDINGS.TreeOfLife-200M-Embeddings
TreeOfLife-200M Embeddings
Pre-computed image embeddings for all images from the TreeOfLife-200M dataset (revision 94bbc0b), sorted by taxonomic hierarchy for efficient filtered access.
This repository hosts embedding configs for TreeOfLife-200M. Each config corresponds to a different embedding model and/or precision. Currently available: BioCLIP 2 (float16) and BioCLIP 2.5 Huge (float16, L2-normalized). Additional configs will be added as new embeddings are generated.
We… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-200M-Embeddings.tech-news-embeddings
Overview
HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023.
To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256.
Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2
Dataset Card for "wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2"
More Information needed
relaion2b-natural-embeddings
LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart)
LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7).
Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings
Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.elite-personas-embeddingsarabic_xvector_embeddings
Arabic Speaker Embeddings extracted from ASC and ClArTTS
There is one speaker embedding for each utterance in the validation set of both datasets. The speaker embeddings are 512-element X-vectors.
Arabic Speech Corpus has 100 files for a single male speaker and ClArTTS has 205 files for a single male speaker.
The X-vectors were extracted using this script, which uses the speechbrain/spkrec-xvect-voxceleb model.
Usage:
from datasets import load_dataset
embeddings_dataset =… See the full description on the dataset page: https://huggingface.co/datasets/herwoww/arabic_xvector_embeddings.datacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
capstone_sakuga_iblip_t5_embeddingssearch-v3-embeddings
Hub Card Search Embeddings (v3)
One-sentence summaries and 1024-d embeddings for 1,173,030 dataset and model cards on the
Hugging Face Hub — 536,870 datasets and 636,160 models. It is the search index behind the revived
librarian-bots/huggingface-semantic-search
backend: you search over a short model-written summary of each card rather than the raw card, and
retrieve against the embedding of that summary.
The cards come from librarian-bots/dataset_cards_with_metadata
and… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/search-v3-embeddings.emilia-yodas-en-speaker-embeddings
Emilia-YODAS English Qwen3-TTS Speaker Embeddings
This dataset contains precomputed speaker embeddings for the English subset of
Emilia-YODAS. Each row maps an Emilia-YODAS sample ID to one speaker embedding
extracted from the corresponding audio.
Dataset Details
Source dataset: amphion/Emilia-Dataset
Source subset: Emilia-YODAS English
Embedding model: Qwen/Qwen3-TTS-12Hz-1.7B-Base
Embedding shape: (2048,)
Embedding dtype: float16
Rows: 4,516,833
Split: train
Additional… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-speaker-embeddings.instagram-political-communication-it-embeddings
Instagram Political Communication (Italy) — Embeddings
This dataset is the companion embeddings dataset ofinstagram-political-communication-it, released as part of the NLP-POL (NLP for Political Communication) project.
It provides vector representations (embeddings) for Instagram posts, comments, sentences, and keyphrases related to the political communication of Italian politicians.
The dataset is designed to support research on:
semantic analysis of political language… See the full description on the dataset page: https://huggingface.co/datasets/NLP-POL/instagram-political-communication-it-embeddings.Tahoe-x1-embeddings
Tahoe-x1 Embeddings on Tahoe-100M
Precomputed embeddings from the Tahoe-x1 foundation model applied to the Tahoe-100M dataset. This dataset provides high-dimensional representations of single-cell transcriptomic profiles from cancer cell lines under small-molecule perturbations.
Overview
This dataset contains cell embeddings generated using the Tahoe-x1-3B model, a 3 billion parameter perturbation-trained single-cell foundation model. The embeddings capture cellular… See the full description on the dataset page: https://huggingface.co/datasets/tahoebio/Tahoe-x1-embeddings.lightonai-embeddings-fine-tuning-reranked-v1
LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2
This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.PaperSeek-OpenAlex-Embeddings
📚 PaperSeek: OpenAlex English Titles & Abstracts (April 2025 Snapshot)
This dataset is part of the PaperSeek framework, a semantic search engine designed for literature discovery using research questions and prior knowledge. PaperSeek is developed as part of a Master's thesis to explore novel approaches in enhancing academic search relevance.
📦 Dataset Overview
Source: OpenAlex
Snapshot Date: April 1st, 2025
Language: English
Contents:
Title
Abstract
Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Grozkal/PaperSeek-OpenAlex-Embeddings.datacomp-small-with-embeddings
Dataset Card for "datacomp-small-with-embeddings"
More Information needed
embeddings_supervised
