CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Qdrant /FineWeb-10B Qdrant-FineWeb-10B Overview Qdrant-FineWeb-10B (Q-FineWeb-10B) is a 10-billion-vector retrieval benchmark derived from FineWeb. Each document is represented with dense and sparse embeddings from Alibaba-NLP/gte-multilingual-base, alongside its original FineWeb payload and metadata. The benchmark also includes exact brute-force ground truth for ~120,000 MS MARCO queries. The dataset includes: 10 billion dense embeddings 10 billion sparse embeddings FineWeb… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/FineWeb-10B.tabular10B<n<100B20 likes39k downloads4d agoHugging Face02max-id /gaianet-qdrant-snapshottext10K<n<100K0 likes5.2k downloads2y agoHugging Face03Qdrant /PubMed-MV PubMed-MultiVector (PubMed-VE) Overview PubMed-VE is a retrieval benchmark built from PubMed abstracts using BGE-M3. Each document is represented in all three formats produced by BGE-M3: ~24 million dense embeddings ~24 million sparse embeddings 8.37B multi-vector (token-level) embeddings (8,370,222,287 exactly; 350.2 tokens/document) Statistic Value Documents 23,898,701 Queries 10,000 Ground truth exact top-1000, one list per modality (dense /… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/PubMed-MV.tabular10M<n<100M4 likes5k downloads24m agoHugging Face04Qdrant /Coyo-VE Coyo-Vector-Embeddings (Coyo-VE) Overview Coyo-VE is a large-scale visual-text-embedding slice of the COYO subset from the LLaVA-OneVision-1.5-Mid-Training-85M dataset, with the original images being sourced from coyo-700m. Each image-caption pair is jointly embedded using Qwen3-VL-Embedding-2B, which accepts mixed image-and-text inputs and encodes their combined visual and textual content into a single 2,048-dimensional dense vector. This produces a unified… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/Coyo-VE.text10M<n<100M2 likes3.5k downloads21d agoHugging Face05Qdrant /arxiv-titles-instructorxl-embeddings arxiv-titles-instructorxl-embeddings This dataset contains 768-dimensional embeddings generated from the arxiv paper titles using InstructorXL model. Each vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The dataset was created using precomputed embeddings exposed by the Alexandria Index. Generation process The embeddings have been generated using the following instruction: Represent the Research Paper title for retrieval;… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings.textsentence-similarity1M<n<10M5 likes3k downloads3y agoHugging Face06Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1536-1M1M OpenAI Embeddings: text-embedding-3-large 1536 dimensions Created: February 2024. Text used for Embedding: title (string) + text (string) Embedding Model: OpenAI text-embedding-3-large This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here textfeature-extraction1M<n<10M14 likes1.3k downloads3y agoHugging Face07Qdrant /dbpedia-entities-openai3-text-embedding-3-large-3072-1M1M OpenAI Embeddings: text-embedding-3-large 3072 dimensions + ada-002 1536 dimensions — parallel dataset Created: February 2024. Text used for Embedding: title (string) + text (string) Embedding Model: text-embedding-3-large This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here textfeature-extraction1M<n<10M27 likes1k downloads3y agoHugging Face08Qdrant /wolt-food-clip-ViT-B-32-embeddings wolt-food-clip-ViT-B-32-embeddings Qdrant's Food Discovery demo relies on the dataset of food images from the Wolt app. Each point in the collection represents a dish with a single image. The image is represented as a vector of 512 float numbers. Generation process The embeddings generated with clip-ViT-B-32 model have been generated using the following code snippet: from PIL import Image from sentence_transformers import SentenceTransformer image_path =… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/wolt-food-clip-ViT-B-32-embeddings.imagefeature-extraction1M<n<10M8 likes452 downloads3y agoHugging Face09Qdrant /arxiv-abstracts-instructorxl-embeddings arxiv-abstracts-instructorxl-embeddings This dataset contains 768-dimensional embeddings generated from the arxiv paper abstracts using InstructorXL model. Each vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The dataset was created using precomputed embeddings exposed by the Alexandria Index. Generation process The embeddings have been generated using the following instruction: Represent the Research Paper abstract for… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-abstracts-instructorxl-embeddings.textsentence-similarity1M<n<10M3 likes443 downloads3y agoHugging Face10Qdrant /hm_ecommerce_products H&M Personalized Fashion Recommendations - Enhanced Dataset Dataset Description This dataset is a processed and enhanced version of the H&M Personalized Fashion Recommendations Kaggle competition dataset. The original dataset has been cleaned and augmented with pre-computed embeddings and accessible image URLs to facilitate fashion recommendation research and multimodal retrieval applications. Dataset Summary The H&M dataset contains rich product metadata… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/hm_ecommerce_products.tabularimage-classification100K<n<1M6 likes242 downloads9mo agoHugging Face11Qdrant /NOAA-Buoy NOAA Buoy meterological data NOAA Buoy Data was downloaded, processed, and cleaned for tasks pertaining to tabular data. The data consists of meteorological measurements. There are two datasets From 1980 through 2022 (denoted with "years" in file names) From Jan 2023 through end of Sept 2023 (denoted with "2023" in file names) The original intended use is for anomaly detection in tabular data. Dataset Details Dataset Description This dataset contains weather… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/NOAA-Buoy.textfeature-extraction100K<n<1M0 likes232 downloads3y agoHugging Face12Qdrant /dbpedia-entities-openai3-text-embedding-3-small-1536-100Ktext100K<n<1M7 likes206 downloads3y agoHugging Face13Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1536-100Ktext100K<n<1M2 likes205 downloads3y agoHugging Face14Qdrant /BGE-m3-1-million-adstabular1M<n<10M1 likes193 downloads5mo agoHugging Face15Qdrant /dbpedia-entities-openai3-text-embedding-3-large-3072-100Ktext100K<n<1M2 likes152 downloads3y agoHugging Face16atitaarora /qdrant_doctextquestion-answeringn<1K0 likes119 downloads2y agoHugging Face17Qdrant /dbpedia-entities-openai3-text-embedding-3-small-512-100Ktext100K<n<1M4 likes70 downloads3y agoHugging Face18atitaarora /qdrant_doc_qnatextn<1K1 likes63 downloads2y agoHugging Face19Qdrant /dbpedia-entities-openai3-text-embedding-3-small-1024-100Ktext100K<n<1M1 likes61 downloads3y agoHugging Face20Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1024-100Ktext100K<n<1M2 likes55 downloads3y agoHugging Face21lukawskikacper /qdrant-landing-page-2024-06-21textn<1K0 likes36 downloads2y agoHugging Face22besartshyti /qdrant_o4_mini_evaltabularn<1K0 likes24 downloads1y agoHugging Face23Qdrant /gte-multilingual-product-ads-1Mtabular1M<n<10M1 likes21 downloads5mo agoHugging Face24Qdrant /gte-multilingual-ads-1Mtabular1M<n<10M1 likes20 downloads5mo agoHugging Face25atitaarora /qdrant_docs_qna_ragastextn<1K0 likes19 downloads3y agoHugging Face26aintech /vdf_qdrant-web-site-docs-2024-04-05This is a dataset created using vector-io text10K<n<100K0 likes19 downloads2y agoHugging Face27omnineura /Qdrantztextn<1K0 likes14 downloads2y agoHugging Face28fastembed /Qdranttextn<1K0 likes8 downloads3y agoHugging Face29amkyawdev /myanmar-v3-qdranttext100K<n<1M0 likes7 downloads3mo agoHugging Face30EmbeddingsOG /propertylm-uk-qdranttabularn<1K0 likes4 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.