CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MongoDB /tech-news-embeddings Overview HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023. To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256. Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.textquestion-answering1M<n<10M6 likes1.6k downloads3y agoHugging Face02flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M12 likes1.4k downloads4y agoHugging Face03flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M8 likes1.1k downloads4y agoHugging Face04flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes780 downloads4y agoHugging Face05Grozkal /PaperSeek-OpenAlex-Embeddings 📚 PaperSeek: OpenAlex English Titles & Abstracts (April 2025 Snapshot) This dataset is part of the PaperSeek framework, a semantic search engine designed for literature discovery using research questions and prior knowledge. PaperSeek is developed as part of a Master's thesis to explore novel approaches in enhancing academic search relevance. 📦 Dataset Overview Source: OpenAlex Snapshot Date: April 1st, 2025 Language: English Contents: Title Abstract Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Grozkal/PaperSeek-OpenAlex-Embeddings.textquestion-answering10M<n<100M5 likes673 downloads1y agoHugging Face06Voxel51 /fiftyone-embeddings-combined FiftyOne Embeddings Dataset This dataset combines the FiftyOne Q&A and function calling datasets with pre-computed embeddings for fast similarity search. Dataset Information Total samples: 28,118 Q&A samples: 14,069 Function samples: 14,049 Embedding model: text-embedding-3-large Embedding dimension: 3072 Schema query: The original question/query text response: The unified response content (either answer text for Q&A or function call text for function… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/fiftyone-embeddings-combined.textquestion-answering10K<n<100K1 likes440 downloads1y agoHugging Face07MongoDB /airbnb_embeddings Overview This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata. It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face. The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.tabularquestion-answering1K<n<10K7 likes412 downloads2y agoHugging Face08flax-sentence-embeddings /stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M20 likes299 downloads4y agoHugging Face09jspringer /open-synthetic-embeddingstextfeature-extraction1M<n<10M3 likes78 downloads1y agoHugging Face10cmpunkmannu /financebench-voyage-finance-2-embeddings FinanceBench (voyage-finance-2 embeddings) Pre-computed embeddings for the FinanceBench corpus. Skips ~$5-15 of Voyage API cost and ~30 minutes of ingest time vs re-embedding from raw PDFs. Intended consumer: the RAG agent at Rishabhmannu/financebench-rag-agent (install: pip install financebench-rag-agent). What's in the box (frozen) Field Value Source corpus FinanceBench (SEC filings: 10-K, 10-Q, 8-K, earnings releases) Parser pypdf (canonical)… See the full description on the dataset page: https://huggingface.co/datasets/cmpunkmannu/financebench-voyage-finance-2-embeddings.tabularsentence-similarity10K<n<100K1 likes35 downloads4mo agoHugging Face11biznetgio /indonesia-law-qa-embeddingstextquestion-answering1K<n<10K6 likes30 downloads2y agoHugging Face12Mercity /ramayana-embeddingsRamayana Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Ramayana, suitable for various NLP tasks, semantic search, and text retrieval purposes. Dataset relesased by Mercity AI! Dataset SourceThe original textual data has been sourced from the repository: Sanskrit Sahitya Data Repository We extend our sincere gratitude to the maintainers of this repository for compiling and sharing valuable Sanskrit literature datasets openly. Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/ramayana-embeddings.textquestion-answering10K<n<100K1 likes27 downloads1y agoHugging Face13BassemE /mulesoft-documentation-embeddings mulesoft-documentation-embeddings MuleSoft Documentation Embeddings for RAG Applications Dataset Information Version: 1.0.0 Created: 2025-09-16T02:41:16.352809 Source: Vector Database License: MIT Language: en Task Categories question-answering, retrieval, knowledge-base Dataset Statistics SkillPilotDataSet_v11 Total Objects: 6430 Unique Properties: 13 Knowledge Sources: mulesoft, user_defined_docs Average Content Length: 5079… See the full description on the dataset page: https://huggingface.co/datasets/BassemE/mulesoft-documentation-embeddings.tabularquestion-answering1K<n<10K0 likes26 downloads1y agoHugging Face14svetfm /epstein-files-nov11-25-house-post-ocr-embeddings Epstein Files Document Embeddings Text embeddings generated from the House Oversight Committee's Epstein document release. Source Dataset This dataset is derived from: tensonaut/EPSTEIN_FILES_20K The source dataset contains OCR'd text from the original House Oversight Committee PDF release. Dataset Structure Field Type Description source_file string Source document filename chunk_index int Position of chunk within document text string Original… See the full description on the dataset page: https://huggingface.co/datasets/svetfm/epstein-files-nov11-25-house-post-ocr-embeddings.texttext-retrieval10K<n<100K3 likes26 downloads10mo agoHugging Face15davenporten /nrc-regulatory-embeddings NRC Regulatory Embeddings 37,734 chunked and embedded NRC nuclear regulatory documents, ready for use in RAG pipelines. Built for the nrc-licensing-rag project, an AI system for analyzing nuclear Combined License Applications (COLAs). Contents Source Documents NUREG-0800 (Standard Review Plan) chapters 1-19 2,436 sections 10 CFR Parts 20, 50, 51, 52, 72, 73, 100 ~504 sections Regulatory Guide Division 1 (1.1-1.262) 242 guides Regulatory Guide Division 4… See the full description on the dataset page: https://huggingface.co/datasets/davenporten/nrc-regulatory-embeddings.texttext-retrieval10K<n<100K2 likes23 downloads6mo agoHugging Face16Laz4rz /wikipedia_stem_small_rag_embeddings STEMWikiSmallRAG with embeddings This dataset contains wikipedia entries from STEM field, unfortunately there is also Business&Economics... but I thought it may contain some useful data as well, even by accident. Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_stem_small_rag_embeddings.texttext-generation100K<n<1M0 likes22 downloads2y agoHugging Face17Mercity /mahabharat-embeddingsMahabharat Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Mahabharat, suitable for various NLP tasks, semantic search, and text retrieval purposes. Dataset relesased by Mercity AI! Dataset SourceThe original textual data has been sourced from the repository: Sanskrit Sahitya Data Repository We extend our sincere gratitude to the maintainers of this repository for compiling and sharing valuable Sanskrit literature datasets openly.… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/mahabharat-embeddings.textquestion-answering10K<n<100K1 likes21 downloads1y agoHugging Face18kevinyulistian /indonesia-law-qa-embeddingstextquestion-answering1K<n<10K0 likes19 downloads4mo agoHugging Face19andfrca /medpt-qwen3-embeddings MedPT with Qwen3 Embeddings About This Derived Dataset This repository is a derived version of the original MedPT dataset, enriched with precomputed semantic embeddings generated from the question field using Qwen3-Embedding-0.6B. The original MedPT corpus, including its medical questions, answers, metadata, collection methodology, curation procedures, annotations, benchmarks, and scientific contributions, was created by Farber et al. (2026). This repository does… See the full description on the dataset page: https://huggingface.co/datasets/andfrca/medpt-qwen3-embeddings.textquestion-answering100K<n<1M0 likes16 downloads1mo agoHugging Face20qarnold /epstein-emails-embeddingstextquestion-answering1K<n<10K0 likes9 downloads10mo agoHugging Face21nairadithya /dictionary-embeddingsEmbeddings generated from the model multi-qa-mpnet-base-dot-v1 being trained on MAKILINGDING/english_dictionary textsentence-similarity100K<n<1M1 likes8 downloads2y agoHugging Face22Mercity /bhagavad_gita-embeddingsBhagavad Gita Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Bhagavad Gita, suitable for various NLP tasks, semantic search, and text retrieval purposes. Dataset relesased by Mercity AI! Dataset SourceThe original textual data has been sourced from the repository: VedaBase.io - Bhagavad Gita Library by ISKCON We extend our sincere gratitude to ISKCON and the maintainers of VedaBase.io for compiling and openly sharing valuable Sanskrit… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/bhagavad_gita-embeddings.textquestion-answeringn<1K1 likes6 downloads1y agoHugging Face23TinaWild09 /epstein-files-nov11-25-house-post-ocr-embeddings Epstein Files Document Embeddings Text embeddings generated from the House Oversight Committee's Epstein document release. Source Dataset This dataset is derived from: tensonaut/EPSTEIN_FILES_20K The source dataset contains OCR'd text from the original House Oversight Committee PDF release. Dataset Structure Field Type Description source_file string Source document filename chunk_index int Position of chunk within document text string Original… See the full description on the dataset page: https://huggingface.co/datasets/TinaWild09/epstein-files-nov11-25-house-post-ocr-embeddings.texttext-retrieval10K<n<100K1 likes6 downloads9mo agoHugging Face24Maki-99 /airbnb_embeddings Overview This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata. It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face. The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/Maki-99/airbnb_embeddings.tabularquestion-answering1K<n<10K0 likes4 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.