CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Kandil7 /Athar-Embeddingstabular1M<n<10M0 likes5.8k downloads5mo agoHugging Face02QuechuaBase /asr-ser-quechua-collao-embeddings ASR-SER embeddings for Quechua Collao This repository contains embeddings only. It does not contain raw audio. These embeddings were extracted for ASR-to-SER transfer experiments on the Quechua Collao emotional speech corpus. The release is intended for reproducible feature-extraction experiments and downstream analysis. Dataset contents One PyTorch tensor per utterance stored as an embedding file under embeddings/ A sanitized metadata table describing utterance… See the full description on the dataset page: https://huggingface.co/datasets/QuechuaBase/asr-ser-quechua-collao-embeddings.tabularaudio-classificationn<1K0 likes1.3k downloads2mo agoHugging Face03flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes885 downloads5y agoHugging Face04coldchair16 /CPRet-Embeddings CPRet-Embeddings This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server. You can explore the retrieval server via the online demo at https://cpret.online/. 📦 Files probs_2609.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-Embeddings.text100K<n<1M1 likes525 downloads18d agoHugging Face05MongoDB /airbnb_embeddings Overview This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata. It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face. The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.tabularquestion-answering1K<n<10K7 likes401 downloads2y agoHugging Face06LeeHarrold /gemma-2b-dictionary-embeddings-all-layers Gemma-2B Dictionary Embeddings - All Layers This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers. Dataset Structure metadata.json: Contains dataset metadata (model info, dimensions, word count) embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26) Usage import pickle from huggingface_hub import hf_hub_download # Download a specific layer layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.tabularn<1K0 likes289 downloads1y agoHugging Face07upctanker /CPRet-Embeddings CPRet-Embeddings This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server. You can explore the retrieval server via the online demo at https://cpret.online/. 📦 Files probs_2606.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/upctanker/CPRet-Embeddings.text100K<n<1M0 likes277 downloads3mo agoHugging Face08prometheus04 /canva-visual-search-embeddings Visual Search Embedding Benchmark: Extending Canva's DINOv2 Evaluation Executive Summary This benchmark extends Canva's January 2025 engineering evaluation which chose DINOv2 for production image replacement. We test three newer models released since then against DINOv2 on 500 design-domain images (advertising posters from the CGL-Dataset). Key Findings Metric Winner Score vs DINOv2 Recall@1 facebook/dinov2-base 1.0000 — Recall@5… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/canva-visual-search-embeddings.textn<1K0 likes200 downloads5mo agoHugging Face09OpenShape /openshape-objaverse-embeddingstextn<1K1 likes199 downloads3y agoHugging Face10mschonhardt /ETP-Eval26-embeddings ETP-Eval26 Embeddings Vorberechnete Vektoren zum Datensatz ETP-Eval26. Inhalt Eine .npz-Datei je Modell, dazu meta.json mit den Zeilenbeschriftungen. Datei Modell Dimension bge-m3.npz BAAI/bge-m3 1024 mE5-large.npz intfloat/multilingual-e5-large 1024 labse.npz sentence-transformers/LaBSE 768 sphilberta.npz bowphs/SPhilBerta 768 qwen3-emb-0.6b.npz Qwen/Qwen3-Embedding-0.6B 1024 qwen3-emb-4b.npz Qwen/Qwen3-Embedding-4B 2560 xlmr.npz… See the full description on the dataset page: https://huggingface.co/datasets/mschonhardt/ETP-Eval26-embeddings.textn<1K0 likes177 downloads2mo agoHugging Face11db-d2 /primevul-codebert-embeddings PrimeVul Embeddings for PU Learning Pre-extracted [CLS] token embeddings from two code models for all functions in the PrimeVul v0.1 vulnerability detection dataset, plus the raw PrimeVul v0.1 JSONL source files. CodeBERT Embeddings (root .npz files) Each .npz file contains frozen CodeBERT embeddings (768-dimensional vectors) for C/C++ functions, along with their labels and CWE type annotations. These were extracted once using a frozen CodeBERT model and are used for… See the full description on the dataset page: https://huggingface.co/datasets/db-d2/primevul-codebert-embeddings.tabulartext-classification100K<n<1M0 likes149 downloads6mo agoHugging Face12DoyyingFace /github-embeddings-doytext1K<n<10K0 likes145 downloads5y agoHugging Face13Joinn /Embeddingstabular100K<n<1M0 likes129 downloads1y agoHugging Face14samwaugh /artefact-embeddings ArteFact Embeddings Dataset A comprehensive collection of vector embeddings for art historical texts, generated using both standard CLIP and specialized PaintingCLIP models to enable semantic search and cross-modal understanding between visual art and textual scholarship. Dataset Overview This dataset contains high-dimensional vector representations of sentences from the ArteFact art historical corpus, enabling semantic search, similarity analysis, and AI-powered research… See the full description on the dataset page: https://huggingface.co/datasets/samwaugh/artefact-embeddings.text1M<n<10M0 likes93 downloads1y agoHugging Face15jspringer /open-synthetic-embeddingstextfeature-extraction1M<n<10M3 likes92 downloads1y agoHugging Face16jumafernandez /d2f-turn-embeddings-soda Turn embeddings for Soda (Dialog2Flow encoder) One 768-d float16 vector per utterance of allenai/soda, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files soda_e_t.f16.npy — numpy array (n_turns, 768), float16;… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-soda.tabularn<1K0 likes86 downloads2mo agoHugging Face17jumafernandez /d2f-turn-embeddings-taskmaster Turn embeddings for Taskmaster (Dialog2Flow encoder) One 768-d float16 vector per utterance of Taskmaster-1/2/3 (Google Research), computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files taskmaster_e_t.f16.npy —… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-taskmaster.tabularn<1K0 likes74 downloads2mo agoHugging Face18matthewcox /paragru-2-4-8m-embeddingstabularn<1K0 likes70 downloads2mo agoHugging Face19jumafernandez /d2f-turn-embeddings-ultrachat Turn embeddings for Ultrachat (Dialog2Flow encoder) One 768-d float16 vector per utterance of openbmb/UltraChat, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files ultrachat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-ultrachat.tabularn<1K0 likes65 downloads2mo agoHugging Face20jordand /echo-embeddings-vctk-tar VCTK Speaker Embeddings (tarred) Items: 109 This dataset ships as a single tar at the repo root. Members preserve paths like VCTK/<id>/audio.mp3 and VCTK/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the CSTR VCTK Corpus. Distributed under CC BY 4.0; attribution required. textn<1K0 likes59 downloads10mo agoHugging Face21jordand /echo-embeddings-expresso-tar Expresso Speaker Embeddings (tarred) Items: 17 This dataset ships as a single tar at the repo root. Members preserve paths like Expresso/<id>/audio.mp3 and Expresso/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the Expresso dataset (INTERSPEECH 2023). Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted. textn<1K0 likes58 downloads10mo agoHugging Face22jordand /echo-embeddings-custom Custom Speaker Embeddings Contains speaker folders within HF-Custom, each with: a precomputed speaker embedding (speaker_latent.safetensors) its corresponding audio (audio.mp3) a metadata file describing the voice and licensing (metadata.json) Licensing:There is no single license for this dataset. Each voice has its own terms stored in its metadata.json. You must check the metadata for any voice you use. audion<1K3 likes58 downloads10mo agoHugging Face23jumafernandez /d2f-turn-embeddings-wildchat Turn embeddings for Wildchat (Dialog2Flow encoder) One 768-d float16 vector per utterance of allenai/WildChat-1M, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files wildchat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-wildchat.tabularn<1K0 likes58 downloads2mo agoHugging Face24jordand /echo-embeddings-ears-tar EARS Speaker Embeddings (tarred) Items: 2568 This dataset ships as a single tar at the repo root. Members preserve paths like EARS/<id>/audio.mp3 and EARS/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the EARS dataset. Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted. text1K<n<10K1 likes52 downloads10mo agoHugging Face25jumafernandez /d2f-turn-embeddings-personachat Turn embeddings for Personachat (Dialog2Flow encoder) One 768-d float16 vector per utterance of Persona-Chat (Zhang et al., 2018), released within ParlAI, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-personachat.tabularn<1K0 likes52 downloads2mo agoHugging Face26Hkang /terminal-bench-4-qwen3-8b-embeddings-train20-seed42-successes Terminal trajectory embedding task subset 20 selected tasks, 1,281 trajectories, and 123,849 state/action pairs. Only trajectories with numeric reward == 1 are retained. The saved vectors are exact selected rows of the existing embeddings; the encoder was not rerun. train/metadata.json matches both tensor row orders. source_row_indices.json records the original row indices; selection.json records selection parameters, source checksums and output checksums. selected_tasks.json… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/terminal-bench-4-qwen3-8b-embeddings-train20-seed42-successes.tabularn<1K0 likes52 downloads4d agoHugging Face27MAMAMARIUS /high_quality_images_embeddingstext1M<n<10M0 likes48 downloads11mo agoHugging Face28totalorganfailure /sec-embeddings-sp500 SEC S&P 500 10q Embeddings Dataset Overview This dataset contains vector embeddings of SEC 10-Q filings for S&P 500 companies Dataset Details Total chunks: 36,927 Embedding model: BAAI/bge-large-en-v1.5 Embedding dimension: 1024 Processing date: 2025-07-15T19:26:08.380350 Quality Metrics Average chunk length: 3182 characters Financial relevance score: 0.146 tabular10K<n<100K0 likes46 downloads1y agoHugging Face29darredondort /decidim-barcelona-proposals-embeddings-768d Decidim Barcelona Proposal Topics 2016-2024 📊 Exploring the top 20 emerging topics from 31,775 citizen proposals in decidim.barcelona, with topic modelling (BERTopic) and deicdim-based open data. 31,775 proposal descriptions from decidim.barcelona (2016-2024), iterating through various parameters and data cleaning techniques, to extract 20 clearly recurrent topics emerging across 270 participatory processes. Sentence embeddings generated using the HuggingFace sentence-transformers… See the full description on the dataset page: https://huggingface.co/datasets/darredondort/decidim-barcelona-proposals-embeddings-768d.tabularsentence-similarity10K<n<100K0 likes45 downloads9mo agoHugging Face30Daniel192341 /RAG-embeddings-storetext10K<n<100K0 likes44 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.