CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hysts-bot-data /daily-papers-embeddingstext10K<n<100K8 likes13k downloads13m agoHugging Face02Major-TOM /Core-AlphaEarth-Embeddings Major TOM Core AlphaEarth Embeddings Subset This is a prototype dataset. It only includes some of the AlphaEarth embeddings stored in Major TOM grid cells. This dataset is mostly aimed at experimentation and prototyping. It is particularly useful to use it along other datasets published within the Major TOM project. Content Field Type Description grid_cell string Major TOM cell year int year of the sample thumbnail image 3-dimensional PCA… See the full description on the dataset page: https://huggingface.co/datasets/Major-TOM/Core-AlphaEarth-Embeddings.image10K<n<100K33 likes11k downloads1y agoHugging Face03lightonai /multilingual-embeddings-pre-training-curated 📚 Collection | 📝 Multilingual Blog | 📝 English Blog Contrastive Multilingual Pre-Training 2.16B query–document pairs across eight languages, plus cross-lingual pairs mDenseOn | mLateOn | DenseOn | LateOn | PyLate | FastPlaid 🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.texttext-retrieval1B<n<10B5 likes9.4k downloads2mo agoHugging Face04justicedao /Caselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.text10M<n<100M0 likes9.1k downloads1y agoHugging Face05scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.5k downloads7mo agoHugging Face06chriswolfram /embeddingstext100K<n<1M0 likes3.4k downloads1y agoHugging Face07laion /Caselaw_Access_Project_embeddingsOriginal Repository: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/ This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.textfeature-extraction10M<n<100M2 likes3.4k downloads1y agoHugging Face08Qdrant /arxiv-titles-instructorxl-embeddings arxiv-titles-instructorxl-embeddings This dataset contains 768-dimensional embeddings generated from the arxiv paper titles using InstructorXL model. Each vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The dataset was created using precomputed embeddings exposed by the Alexandria Index. Generation process The embeddings have been generated using the following instruction: Represent the Research Paper title for retrieval;… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings.textsentence-similarity1M<n<10M5 likes3.2k downloads3y agoHugging Face09kshitijd /platonic-embeddingstimeseries1M<n<10M0 likes2.5k downloads5mo agoHugging Face10GlobalCampus /openalex-multilingual-embeddings OpenAlex Multilingual Embeddings This dataset contains multilingual text embeddings of all records in OpenAlex with a title or an abstract from the snapshot of 2023-10-20. The dataset was created for the FORAS project to investigate the efficacy of different methods of searching in databases of academic publications. All scripts will be available in a GitHub repository. The project is supported by a grant from the Dutch Research Council (grant no. 406.22.GO.048)… See the full description on the dataset page: https://huggingface.co/datasets/GlobalCampus/openalex-multilingual-embeddings.text100M<n<1B0 likes2.3k downloads3y agoHugging Face11keypa /vision-adapter-embeddings Vision Adapter MoonViT Embeddings Precomputed visual embeddings used to train lightweight vision→LLM projectors without re-running a vision tower: each row is the frozen MoonViT-V2 output for one training image, stored as raw bfloat16 bytes. Shards: 103 Parquet files (data/emb_0000.parquet … data/emb_0102.parquet), 1360 rows each, ~139k rows total, ~1.9 TB. Schema per row: column type meaning key string embedding id, embeddings/<sha1[:20]>.pt; matches emb in… See the full description on the dataset page: https://huggingface.co/datasets/keypa/vision-adapter-embeddings.textimage-to-text100K<n<1M0 likes2.2k downloads2d agoHugging Face12lightonai /embeddings-fine-tuning Overview This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version. This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning.text10M<n<100M24 likes2.1k downloads3mo agoHugging Face13Marqo /benchmark-embeddings Marqo Benchmark Embeddings This dataset contains a large collection of embeddings from popular models on benchmark datasets. In addition to this, the passage data also includes the local intrinsic dimensionality (LID) for every vector considering its exact nearest 100 neighbours, LID is calculated using a Maximum Likelihood Estimation based approach. Below is a list of the datasets and the models, every datasets queries and passages are embedded with every model.… See the full description on the dataset page: https://huggingface.co/datasets/Marqo/benchmark-embeddings.text100M<n<1B6 likes2k downloads2y agoHugging Face14DSIMB /PATHOS-PLM-EMBEDDINGS PATHOS PLM Embeddings Precomputed protein language model (PLM) embeddings for missense substitutions and wild-type residues in 20,416 human SwissProt proteins. These embeddings are used by PATHOS to predict the pathogenicity of missense mutations. Paper: http://dx.doi.org/10.1016/j.ailsci.2026.100165 Dataset Structure The repository contains two config families for each PLM: Mutation configs: <model> stores embeddings for generated missense substitutions. Wild-type… See the full description on the dataset page: https://huggingface.co/datasets/DSIMB/PATHOS-PLM-EMBEDDINGS.textfeature-extraction100M<n<1B3 likes1.9k downloads4mo agoHugging Face15imageomics /TreeOfLife-200M-Embeddings TreeOfLife-200M Embeddings Pre-computed image embeddings for all images from the TreeOfLife-200M dataset (revision 94bbc0b), sorted by taxonomic hierarchy for efficient filtered access. This repository hosts embedding configs for TreeOfLife-200M. Each config corresponds to a different embedding model and/or precision. Currently available: BioCLIP 2 (float16) and BioCLIP 2.5 Huge (float16, L2-normalized). Additional configs will be added as new embeddings are generated. We… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-200M-Embeddings.imagefeature-extraction100M<n<1B1 likes1.8k downloads29d agoHugging Face16MongoDB /tech-news-embeddings Overview HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023. To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256. Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.textquestion-answering1M<n<10M6 likes1.6k downloads3y agoHugging Face17maloyan /wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2 Dataset Card for "wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2" More Information needed tabular10M<n<100M4 likes1.6k downloads3y agoHugging Face18andropar /relaion2b-natural-embeddings LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart) LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7). Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.tabularfeature-extraction100M<n<1B1 likes1.5k downloads6mo agoHugging Face19altaidevorg /elite-personas-embeddingstext100M<n<1B0 likes1.4k downloads10mo agoHugging Face20herwoww /arabic_xvector_embeddings Arabic Speaker Embeddings extracted from ASC and ClArTTS There is one speaker embedding for each utterance in the validation set of both datasets. The speaker embeddings are 512-element X-vectors. Arabic Speech Corpus has 100 files for a single male speaker and ClArTTS has 205 files for a single male speaker. The X-vectors were extracted using this script, which uses the speechbrain/spkrec-xvect-voxceleb model. Usage: from datasets import load_dataset embeddings_dataset =… See the full description on the dataset page: https://huggingface.co/datasets/herwoww/arabic_xvector_embeddings.texttext-to-speechn<1K5 likes1.1k downloads2y agoHugging Face21nielsr /datacomp-small-with-text-embeddings Dataset Card for "datacomp-small-with-text-embeddings" More Information needed image10M<n<100M0 likes1.1k downloads3y agoHugging Face22Hemabhushan /capstone_sakuga_iblip_t5_embeddingstabular10K<n<100K0 likes1k downloads2y agoHugging Face23davanstrien /search-v3-embeddings Hub Card Search Embeddings (v3) One-sentence summaries and 1024-d embeddings for 1,173,030 dataset and model cards on the Hugging Face Hub — 536,870 datasets and 636,160 models. It is the search index behind the revived librarian-bots/huggingface-semantic-search backend: you search over a short model-written summary of each card rather than the raw card, and retrieve against the embedding of that summary. The cards come from librarian-bots/dataset_cards_with_metadata and… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/search-v3-embeddings.tabularfeature-extraction1M<n<10M4 likes1k downloads1mo agoHugging Face24duplexio /emilia-yodas-en-speaker-embeddings Emilia-YODAS English Qwen3-TTS Speaker Embeddings This dataset contains precomputed speaker embeddings for the English subset of Emilia-YODAS. Each row maps an Emilia-YODAS sample ID to one speaker embedding extracted from the corresponding audio. Dataset Details Source dataset: amphion/Emilia-Dataset Source subset: Emilia-YODAS English Embedding model: Qwen/Qwen3-TTS-12Hz-1.7B-Base Embedding shape: (2048,) Embedding dtype: float16 Rows: 4,516,833 Split: train Additional… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-speaker-embeddings.textfeature-extraction1M<n<10M0 likes866 downloads4mo agoHugging Face25NLP-POL /instagram-political-communication-it-embeddings Instagram Political Communication (Italy) — Embeddings This dataset is the companion embeddings dataset ofinstagram-political-communication-it, released as part of the NLP-POL (NLP for Political Communication) project. It provides vector representations (embeddings) for Instagram posts, comments, sentences, and keyphrases related to the political communication of Italian politicians. The dataset is designed to support research on: semantic analysis of political language… See the full description on the dataset page: https://huggingface.co/datasets/NLP-POL/instagram-political-communication-it-embeddings.textfeature-extraction1M<n<10M1 likes810 downloads9mo agoHugging Face26tahoebio /Tahoe-x1-embeddings Tahoe-x1 Embeddings on Tahoe-100M Precomputed embeddings from the Tahoe-x1 foundation model applied to the Tahoe-100M dataset. This dataset provides high-dimensional representations of single-cell transcriptomic profiles from cancer cell lines under small-molecule perturbations. Overview This dataset contains cell embeddings generated using the Tahoe-x1-3B model, a 3 billion parameter perturbation-trained single-cell foundation model. The embeddings capture cellular… See the full description on the dataset page: https://huggingface.co/datasets/tahoebio/Tahoe-x1-embeddings.textfeature-extraction10M<n<100M6 likes775 downloads10mo agoHugging Face27cross-encoder /lightonai-embeddings-fine-tuning-reranked-v1 LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2 This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.texttext-ranking10M<n<100M10 likes771 downloads4mo agoHugging Face28Grozkal /PaperSeek-OpenAlex-Embeddings 📚 PaperSeek: OpenAlex English Titles & Abstracts (April 2025 Snapshot) This dataset is part of the PaperSeek framework, a semantic search engine designed for literature discovery using research questions and prior knowledge. PaperSeek is developed as part of a Master's thesis to explore novel approaches in enhancing academic search relevance. 📦 Dataset Overview Source: OpenAlex Snapshot Date: April 1st, 2025 Language: English Contents: Title Abstract Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Grozkal/PaperSeek-OpenAlex-Embeddings.textquestion-answering10M<n<100M5 likes673 downloads1y agoHugging Face29nielsr /datacomp-small-with-embeddings Dataset Card for "datacomp-small-with-embeddings" More Information needed image10M<n<100M0 likes654 downloads3y agoHugging Face30lightonai /embeddings_supervisedtext1M<n<10M13 likes546 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.