datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Athar-Embeddingsasr-ser-quechua-collao-embeddings
ASR-SER embeddings for Quechua Collao
This repository contains embeddings only. It does not contain raw audio.
These embeddings were extracted for ASR-to-SER transfer experiments on the Quechua Collao emotional speech corpus. The release is intended for reproducible feature-extraction experiments and downstream analysis.
Dataset contents
One PyTorch tensor per utterance stored as an embedding file under embeddings/
A sanitized metadata table describing utterance… See the full description on the dataset page: https://huggingface.co/datasets/QuechuaBase/asr-ser-quechua-collao-embeddings.stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml
Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]}
The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0
If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file.
This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.CPRet-Embeddings
CPRet-Embeddings
This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server.
You can explore the retrieval server via the online demo at https://cpret.online/.
📦 Files
probs_2609.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-Embeddings.airbnb_embeddings
Overview
This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata.
It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face.
The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.gemma-2b-dictionary-embeddings-all-layers
Gemma-2B Dictionary Embeddings - All Layers
This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers.
Dataset Structure
metadata.json: Contains dataset metadata (model info, dimensions, word count)
embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26)
Usage
import pickle
from huggingface_hub import hf_hub_download
# Download a specific layer
layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.CPRet-Embeddings
CPRet-Embeddings
This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server.
You can explore the retrieval server via the online demo at https://cpret.online/.
📦 Files
probs_2606.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/upctanker/CPRet-Embeddings.ETP-Eval26-embeddings
ETP-Eval26 Embeddings
Vorberechnete Vektoren zum Datensatz
ETP-Eval26.
Inhalt
Eine .npz-Datei je Modell, dazu meta.json mit den Zeilenbeschriftungen.
Datei
Modell
Dimension
bge-m3.npz
BAAI/bge-m3
1024
mE5-large.npz
intfloat/multilingual-e5-large
1024
labse.npz
sentence-transformers/LaBSE
768
sphilberta.npz
bowphs/SPhilBerta
768
qwen3-emb-0.6b.npz
Qwen/Qwen3-Embedding-0.6B
1024
qwen3-emb-4b.npz
Qwen/Qwen3-Embedding-4B
2560
xlmr.npz… See the full description on the dataset page: https://huggingface.co/datasets/mschonhardt/ETP-Eval26-embeddings.canva-visual-search-embeddings
Visual Search Embedding Benchmark: Extending Canva's DINOv2 Evaluation
Executive Summary
This benchmark extends Canva's January 2025 engineering evaluation
which chose DINOv2 for production image replacement. We test three newer models released since then
against DINOv2 on 500 design-domain images (advertising posters from the CGL-Dataset).
Key Findings
Metric
Winner
Score
vs DINOv2
Recall@1
facebook/dinov2-base
1.0000
—
Recall@5… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/canva-visual-search-embeddings.openshape-objaverse-embeddingsprimevul-codebert-embeddings
PrimeVul Embeddings for PU Learning
Pre-extracted [CLS] token embeddings from two code models for all functions in the PrimeVul v0.1 vulnerability detection dataset, plus the raw PrimeVul v0.1 JSONL source files.
CodeBERT Embeddings (root .npz files)
Each .npz file contains frozen CodeBERT embeddings (768-dimensional vectors) for C/C++ functions, along with their labels and CWE type annotations. These were extracted once using a frozen CodeBERT model and are used for… See the full description on the dataset page: https://huggingface.co/datasets/db-d2/primevul-codebert-embeddings.github-embeddings-doyEmbeddingsd2f-turn-embeddings-soda
Turn embeddings for Soda (Dialog2Flow encoder)
One 768-d float16 vector per utterance of allenai/soda, computed with the frozen
encoder sergioburdisso/dialog2flow-joint-bert-base
(SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder
truncates inputs at 64 tokens. If you use these embeddings, please cite the
Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus.
Files
soda_e_t.f16.npy — numpy array (n_turns, 768), float16;… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-soda.artefact-embeddings
ArteFact Embeddings Dataset
A comprehensive collection of vector embeddings for art historical texts, generated using both standard CLIP and specialized PaintingCLIP models to enable semantic search and cross-modal understanding between visual art and textual scholarship.
Dataset Overview
This dataset contains high-dimensional vector representations of sentences from the ArteFact art historical corpus, enabling semantic search, similarity analysis, and AI-powered research… See the full description on the dataset page: https://huggingface.co/datasets/samwaugh/artefact-embeddings.open-synthetic-embeddingsd2f-turn-embeddings-taskmaster
Turn embeddings for Taskmaster (Dialog2Flow encoder)
One 768-d float16 vector per utterance of Taskmaster-1/2/3 (Google Research), computed with the frozen
encoder sergioburdisso/dialog2flow-joint-bert-base
(SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder
truncates inputs at 64 tokens. If you use these embeddings, please cite the
Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus.
Files
taskmaster_e_t.f16.npy —… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-taskmaster.paragru-2-4-8m-embeddingsvirtual-news-embeddingsd2f-turn-embeddings-ultrachat
Turn embeddings for Ultrachat (Dialog2Flow encoder)
One 768-d float16 vector per utterance of openbmb/UltraChat, computed with the frozen
encoder sergioburdisso/dialog2flow-joint-bert-base
(SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder
truncates inputs at 64 tokens. If you use these embeddings, please cite the
Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus.
Files
ultrachat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-ultrachat.echo-embeddings-custom
Custom Speaker Embeddings
Contains speaker folders within HF-Custom, each with:
a precomputed speaker embedding (speaker_latent.safetensors)
its corresponding audio (audio.mp3)
a metadata file describing the voice and licensing (metadata.json)
Licensing:There is no single license for this dataset. Each voice has its own terms stored
in its metadata.json. You must check the metadata for any voice you use.
sec-embeddings-sp500
SEC S&P 500 10q Embeddings Dataset
Overview
This dataset contains vector embeddings of SEC 10-Q filings for S&P 500 companies
Dataset Details
Total chunks: 36,927
Embedding model: BAAI/bge-large-en-v1.5
Embedding dimension: 1024
Processing date: 2025-07-15T19:26:08.380350
Quality Metrics
Average chunk length: 3182 characters
Financial relevance score: 0.146
d2f-turn-embeddings-wildchat
Turn embeddings for Wildchat (Dialog2Flow encoder)
One 768-d float16 vector per utterance of allenai/WildChat-1M, computed with the frozen
encoder sergioburdisso/dialog2flow-joint-bert-base
(SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder
truncates inputs at 64 tokens. If you use these embeddings, please cite the
Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus.
Files
wildchat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-wildchat.echo-embeddings-vctk-tar
VCTK Speaker Embeddings (tarred)
Items: 109
This dataset ships as a single tar at the repo root. Members preserve paths like
VCTK/<id>/audio.mp3 and VCTK/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the CSTR VCTK Corpus. Distributed under CC BY 4.0; attribution required.
echo-embeddings-expresso-tar
Expresso Speaker Embeddings (tarred)
Items: 17
This dataset ships as a single tar at the repo root. Members preserve paths like
Expresso/<id>/audio.mp3 and Expresso/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the Expresso dataset (INTERSPEECH 2023). Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted.
mr-right-zhtw-embeddings
Mr. Right zh-TW — pre-computed document embeddings
The deployed document bank (zhen_img_v2) for the
Mr. Right zh-TW corpus:
one 4096-dimensional vector per document, covering all 769,245 documents.
Encoding all of them takes multiple GPU-days and requires the full 1.2 TB image tree.
This repository exists so you can skip both and go straight to retrieval.
file
shape / size
notes
emb.npy
(769245, 4096) float16, 6.3 GB
L2-normalised — cosine similarity is a plain dot… See the full description on the dataset page: https://huggingface.co/datasets/ericssonbear/mr-right-zhtw-embeddings.high_quality_images_embeddingsd2f-turn-embeddings-personachat
Turn embeddings for Personachat (Dialog2Flow encoder)
One 768-d float16 vector per utterance of Persona-Chat (Zhang et al., 2018), released within ParlAI, computed with the frozen
encoder sergioburdisso/dialog2flow-joint-bert-base
(SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder
truncates inputs at 64 tokens. If you use these embeddings, please cite the
Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus.
Files… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-personachat.paws-jsonl
Introduction
This dataset is a jsonl format for PAWS dataset from: https://github.com/google-research-datasets/paws. It only contains the PAWS-Wiki Labeled (Final) and
PAWS-Wiki Labeled (Swap-only) training sections of the original PAWS dataset. Duplicates data are removed.
Each line contains a dict in the following format:
{"guid": <id>, "texts": [anchor, positive]} or
{"guid": <id>, "texts": [anchor, positive, negative]}
positives_negatives.jsonl.gz: 24,723… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/paws-jsonl.echo-embeddings-ears-tar
EARS Speaker Embeddings (tarred)
Items: 2568
This dataset ships as a single tar at the repo root. Members preserve paths like
EARS/<id>/audio.mp3 and EARS/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the EARS dataset. Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted.
