CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Kandil7 /Athar-Embeddingstabular1M<n<10M0 likes5.7k downloads5mo agoHugging Face02QuechuaBase /asr-ser-quechua-collao-embeddings ASR-SER embeddings for Quechua Collao This repository contains embeddings only. It does not contain raw audio. These embeddings were extracted for ASR-to-SER transfer experiments on the Quechua Collao emotional speech corpus. The release is intended for reproducible feature-extraction experiments and downstream analysis. Dataset contents One PyTorch tensor per utterance stored as an embedding file under embeddings/ A sanitized metadata table describing utterance… See the full description on the dataset page: https://huggingface.co/datasets/QuechuaBase/asr-ser-quechua-collao-embeddings.tabularaudio-classificationn<1K0 likes1.2k downloads2mo agoHugging Face03flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes856 downloads5y agoHugging Face04coldchair16 /CPRet-Embeddings CPRet-Embeddings This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server. You can explore the retrieval server via the online demo at https://cpret.online/. 📦 Files probs_2609.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-Embeddings.text100K<n<1M1 likes546 downloads17d agoHugging Face05MongoDB /airbnb_embeddings Overview This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata. It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face. The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.tabularquestion-answering1K<n<10K7 likes404 downloads2y agoHugging Face06LeeHarrold /gemma-2b-dictionary-embeddings-all-layers Gemma-2B Dictionary Embeddings - All Layers This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers. Dataset Structure metadata.json: Contains dataset metadata (model info, dimensions, word count) embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26) Usage import pickle from huggingface_hub import hf_hub_download # Download a specific layer layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.tabularn<1K0 likes281 downloads1y agoHugging Face07upctanker /CPRet-Embeddings CPRet-Embeddings This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server. You can explore the retrieval server via the online demo at https://cpret.online/. 📦 Files probs_2606.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/upctanker/CPRet-Embeddings.text100K<n<1M0 likes278 downloads3mo agoHugging Face08mschonhardt /ETP-Eval26-embeddings ETP-Eval26 Embeddings Vorberechnete Vektoren zum Datensatz ETP-Eval26. Inhalt Eine .npz-Datei je Modell, dazu meta.json mit den Zeilenbeschriftungen. Datei Modell Dimension bge-m3.npz BAAI/bge-m3 1024 mE5-large.npz intfloat/multilingual-e5-large 1024 labse.npz sentence-transformers/LaBSE 768 sphilberta.npz bowphs/SPhilBerta 768 qwen3-emb-0.6b.npz Qwen/Qwen3-Embedding-0.6B 1024 qwen3-emb-4b.npz Qwen/Qwen3-Embedding-4B 2560 xlmr.npz… See the full description on the dataset page: https://huggingface.co/datasets/mschonhardt/ETP-Eval26-embeddings.textn<1K0 likes204 downloads2mo agoHugging Face09prometheus04 /canva-visual-search-embeddings Visual Search Embedding Benchmark: Extending Canva's DINOv2 Evaluation Executive Summary This benchmark extends Canva's January 2025 engineering evaluation which chose DINOv2 for production image replacement. We test three newer models released since then against DINOv2 on 500 design-domain images (advertising posters from the CGL-Dataset). Key Findings Metric Winner Score vs DINOv2 Recall@1 facebook/dinov2-base 1.0000 — Recall@5… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/canva-visual-search-embeddings.textn<1K0 likes186 downloads5mo agoHugging Face10OpenShape /openshape-objaverse-embeddingstextn<1K1 likes161 downloads3y agoHugging Face11db-d2 /primevul-codebert-embeddings PrimeVul Embeddings for PU Learning Pre-extracted [CLS] token embeddings from two code models for all functions in the PrimeVul v0.1 vulnerability detection dataset, plus the raw PrimeVul v0.1 JSONL source files. CodeBERT Embeddings (root .npz files) Each .npz file contains frozen CodeBERT embeddings (768-dimensional vectors) for C/C++ functions, along with their labels and CWE type annotations. These were extracted once using a frozen CodeBERT model and are used for… See the full description on the dataset page: https://huggingface.co/datasets/db-d2/primevul-codebert-embeddings.tabulartext-classification100K<n<1M0 likes153 downloads6mo agoHugging Face12DoyyingFace /github-embeddings-doytext1K<n<10K0 likes145 downloads5y agoHugging Face13Joinn /Embeddingstabular100K<n<1M0 likes129 downloads1y agoHugging Face14jumafernandez /d2f-turn-embeddings-soda Turn embeddings for Soda (Dialog2Flow encoder) One 768-d float16 vector per utterance of allenai/soda, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files soda_e_t.f16.npy — numpy array (n_turns, 768), float16;… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-soda.tabularn<1K0 likes103 downloads2mo agoHugging Face15samwaugh /artefact-embeddings ArteFact Embeddings Dataset A comprehensive collection of vector embeddings for art historical texts, generated using both standard CLIP and specialized PaintingCLIP models to enable semantic search and cross-modal understanding between visual art and textual scholarship. Dataset Overview This dataset contains high-dimensional vector representations of sentences from the ArteFact art historical corpus, enabling semantic search, similarity analysis, and AI-powered research… See the full description on the dataset page: https://huggingface.co/datasets/samwaugh/artefact-embeddings.text1M<n<10M0 likes93 downloads1y agoHugging Face16jspringer /open-synthetic-embeddingstextfeature-extraction1M<n<10M3 likes92 downloads1y agoHugging Face17jumafernandez /d2f-turn-embeddings-taskmaster Turn embeddings for Taskmaster (Dialog2Flow encoder) One 768-d float16 vector per utterance of Taskmaster-1/2/3 (Google Research), computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files taskmaster_e_t.f16.npy —… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-taskmaster.tabularn<1K0 likes71 downloads2mo agoHugging Face18matthewcox /paragru-2-4-8m-embeddingstabularn<1K0 likes70 downloads2mo agoHugging Face19Hiba03 /virtual-news-embeddingstext10K<n<100K0 likes64 downloads5mo agoHugging Face20jumafernandez /d2f-turn-embeddings-ultrachat Turn embeddings for Ultrachat (Dialog2Flow encoder) One 768-d float16 vector per utterance of openbmb/UltraChat, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files ultrachat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-ultrachat.tabularn<1K0 likes63 downloads2mo agoHugging Face21jordand /echo-embeddings-custom Custom Speaker Embeddings Contains speaker folders within HF-Custom, each with: a precomputed speaker embedding (speaker_latent.safetensors) its corresponding audio (audio.mp3) a metadata file describing the voice and licensing (metadata.json) Licensing:There is no single license for this dataset. Each voice has its own terms stored in its metadata.json. You must check the metadata for any voice you use. audion<1K3 likes58 downloads10mo agoHugging Face22totalorganfailure /sec-embeddings-sp500 SEC S&P 500 10q Embeddings Dataset Overview This dataset contains vector embeddings of SEC 10-Q filings for S&P 500 companies Dataset Details Total chunks: 36,927 Embedding model: BAAI/bge-large-en-v1.5 Embedding dimension: 1024 Processing date: 2025-07-15T19:26:08.380350 Quality Metrics Average chunk length: 3182 characters Financial relevance score: 0.146 tabular10K<n<100K0 likes57 downloads1y agoHugging Face23jumafernandez /d2f-turn-embeddings-wildchat Turn embeddings for Wildchat (Dialog2Flow encoder) One 768-d float16 vector per utterance of allenai/WildChat-1M, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files wildchat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-wildchat.tabularn<1K0 likes57 downloads2mo agoHugging Face24jordand /echo-embeddings-vctk-tar VCTK Speaker Embeddings (tarred) Items: 109 This dataset ships as a single tar at the repo root. Members preserve paths like VCTK/<id>/audio.mp3 and VCTK/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the CSTR VCTK Corpus. Distributed under CC BY 4.0; attribution required. textn<1K0 likes56 downloads10mo agoHugging Face25jordand /echo-embeddings-expresso-tar Expresso Speaker Embeddings (tarred) Items: 17 This dataset ships as a single tar at the repo root. Members preserve paths like Expresso/<id>/audio.mp3 and Expresso/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the Expresso dataset (INTERSPEECH 2023). Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted. textn<1K0 likes55 downloads10mo agoHugging Face26ericssonbear /mr-right-zhtw-embeddings Mr. Right zh-TW — pre-computed document embeddings The deployed document bank (zhen_img_v2) for the Mr. Right zh-TW corpus: one 4096-dimensional vector per document, covering all 769,245 documents. Encoding all of them takes multiple GPU-days and requires the full 1.2 TB image tree. This repository exists so you can skip both and go straight to retrieval. file shape / size notes emb.npy (769245, 4096) float16, 6.3 GB L2-normalised — cosine similarity is a plain dot… See the full description on the dataset page: https://huggingface.co/datasets/ericssonbear/mr-right-zhtw-embeddings.text-retrieval100K<n<1M0 likes54 downloads1mo agoHugging Face27MAMAMARIUS /high_quality_images_embeddingstext1M<n<10M0 likes52 downloads11mo agoHugging Face28jumafernandez /d2f-turn-embeddings-personachat Turn embeddings for Personachat (Dialog2Flow encoder) One 768-d float16 vector per utterance of Persona-Chat (Zhang et al., 2018), released within ParlAI, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-personachat.tabularn<1K0 likes48 downloads2mo agoHugging Face29flax-sentence-embeddings /paws-jsonl Introduction This dataset is a jsonl format for PAWS dataset from: https://github.com/google-research-datasets/paws. It only contains the PAWS-Wiki Labeled (Final) and PAWS-Wiki Labeled (Swap-only) training sections of the original PAWS dataset. Duplicates data are removed. Each line contains a dict in the following format: {"guid": <id>, "texts": [anchor, positive]} or {"guid": <id>, "texts": [anchor, positive, negative]} positives_negatives.jsonl.gz: 24,723… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/paws-jsonl.text10K<n<100K1 likes47 downloads5y agoHugging Face30jordand /echo-embeddings-ears-tar EARS Speaker Embeddings (tarred) Items: 2568 This dataset ships as a single tar at the repo root. Members preserve paths like EARS/<id>/audio.mp3 and EARS/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the EARS dataset. Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted. text1K<n<10K1 likes46 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.