CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hotchpotch /bekko-embedding-v1-unsupervised Bekko Embedding v1 Unsupervised Training Data This is the unsupervised training dataset used for the bekko-embedding-v1 embedding model family. It contains pair and triplet examples for embedding pretraining, published as Hugging Face dataset subsets so that each source/subset/task combination can be loaded independently. For the full training recipe and technical details, see Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/bekko-embedding-v1-unsupervised.text-retrieval4 likes29k downloads2mo agoHugging Face02thanminh01 /WSI_Embedding WSI_Embedding Patch and slide embeddings for whole-slide images, plus the run records used to produce them. Original WSIs are not in this repo. They stay on HPC / the source dataset remotes. Do not expect .svs / .tif here. Repo: thanminh01/WSI_Embedding Top-level layout Path What it is Download when embeddings/<dataset>/<model>/<mag>x_<patch>px_<overlap>px_overlap/ Encoder tensors (.h5 / .pt / WSI-LLaVA json) and per-model TRIDENT _config_ / _logs_ You… See the full description on the dataset page: https://huggingface.co/datasets/thanminh01/WSI_Embedding.imagefeature-extraction1K<n<10K0 likes26k downloads7d agoHugging Face03asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes21k downloads2y agoHugging Face04asahi417 /seamless-align-enA-esA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes20k downloads2y agoHugging Face05VoCuc /vlm-teacher-embedding0 likes19k downloads16h agoHugging Face06asahi417 /seamless-align-enA-frA.speaker-embedding.hubert-xltabular1M<n<10M0 likes16k downloads2y agoHugging Face07asahi417 /seamless-align-enA-jaA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes16k downloads2y agoHugging Face08hysts-bot-data /daily-papers-embeddingstext10K<n<100K8 likes14k downloads6h agoHugging Face09Major-TOM /Core-AlphaEarth-Embeddings Major TOM Core AlphaEarth Embeddings Subset This is a prototype dataset. It only includes some of the AlphaEarth embeddings stored in Major TOM grid cells. This dataset is mostly aimed at experimentation and prototyping. It is particularly useful to use it along other datasets published within the Major TOM project. Content Field Type Description grid_cell string Major TOM cell year int year of the sample thumbnail image 3-dimensional PCA… See the full description on the dataset page: https://huggingface.co/datasets/Major-TOM/Core-AlphaEarth-Embeddings.image10K<n<100K33 likes11k downloads1y agoHugging Face10asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes11k downloads2y agoHugging Face11asahi417 /seamless-align-enA-hiA.speaker-embedding.hubert-xltabular100K<n<1M0 likes10k downloads2y agoHugging Face12asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes10k downloads2y agoHugging Face13asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.9k downloads2y agoHugging Face14asahi417 /seamless-align-enA-zhA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes9.7k downloads2y agoHugging Face15shash42 /forecast-news-embeddings Forecast News Embeddings Precomputed LanceDB table used for keyword, semantic, and hybrid retrieval in forecast-sim and future-sim. Snapshot 7,911,857 indexed source articles 16,207,764 text chunks Coverage: 2023-01-11 through 2026-08-31 Snapshot published: 2026-09-18 Lance dataset version: 856 Total artifact size: approximately 303.2 GiB Articles with empty searchable text are not represented. Long articles can produce multiple chunks, so the chunk count is… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news-embeddings.2 likes9.7k downloads3d agoHugging Face16LetsChurch /bible-embeddings Bible Embeddings A comprehensive tool for generating and evaluating Bible verse embeddings using various state-of-the-art embedding models. This project supports both commercial APIs (OpenAI, Google Gemini, Voyage AI) and open-source models (HuggingFace sentence-transformers) for semantic search across biblical texts. Setup This project is managed with uv. Make sure you have uv installed, then set up the project: # Install dependencies uv sync # Install specific… See the full description on the dataset page: https://huggingface.co/datasets/LetsChurch/bible-embeddings.5 likes9.4k downloads6mo agoHugging Face17sebasmos /latent-sr-embeddings Latent-SR Embeddings: Precomputed VAE Latents for Medical Image Super-Resolution Precomputed VAE latent embeddings from the paper: "Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution"Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu, Sahil Kapadia, Hillary Clinton Kasimbazi, Leo Kinyera, Emmanuel Paul Kwesiga, Sri Sri Jaithra Varma Manthena, Luis Filipe Nakayama, Ninsiima Doreen, Leo Anthony Celi.arXiv:2604.12152 (2026)… See the full description on the dataset page: https://huggingface.co/datasets/sebasmos/latent-sr-embeddings.textimage-to-image10K<n<100K1 likes9.1k downloads3mo agoHugging Face18justicedao /Caselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.text10M<n<100M0 likes9.1k downloads1y agoHugging Face19asahi417 /seamless-align-enA-frA.speaker-embedding.w2vbert-600mtabular1M<n<10M0 likes8.9k downloads2y agoHugging Face20lightonai /multilingual-embeddings-pre-training-curated 📚 Collection | 📝 Multilingual Blog | 📝 English Blog Contrastive Multilingual Pre-Training 2.16B query–document pairs across eight languages, plus cross-lingual pairs mDenseOn | mLateOn | DenseOn | LateOn | PyLate | FastPlaid 🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.texttext-retrieval1B<n<10B5 likes8.3k downloads2mo agoHugging Face21asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes8.1k downloads2y agoHugging Face22asahi417 /seamless-align-enA-zhA.speaker-embedding.hubert-xltabular100K<n<1M0 likes7k downloads2y agoHugging Face23asahi417 /seamless-align-enA-koA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.9k downloads2y agoHugging Face24lightonai /embeddings-pre-training Overview This large-scale dataset is designed for pre-training state-of-the-art text embedding models. Its goal is to reproduce and build upon the data recipe described in the mGTE technical report (Zhang et al., 2024), which details the data sources used to train the GTE family of embedding models but does not release the data itself. We assembled this dataset as part of a research effort to understand how data composition affects retrieval model quality. Our experiments… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training.51 likes6.9k downloads26d agoHugging Face25scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.5k downloads6mo agoHugging Face26asahi417 /seamless-align-enA-hiA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.5k downloads2y agoHugging Face27asahi417 /seamless-align-enA-viA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.2k downloads2y agoHugging Face28asahi417 /seamless-align-enA-jaA.speaker-embedding.hubert-xltabular100K<n<1M0 likes6k downloads2y agoHugging Face29asahi417 /seamless-align-enA-hiA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.8k downloads2y agoHugging Face30asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.7k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.