CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MongoDB /tech-news-embeddings Overview HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023. To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256. Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.textquestion-answering1M<n<10M6 likes1.8k downloads3y agoHugging Face02flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M12 likes1.5k downloads4y agoHugging Face03flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M8 likes1.1k downloads4y agoHugging Face04flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes737 downloads4y agoHugging Face05Grozkal /PaperSeek-OpenAlex-Embeddings 📚 PaperSeek: OpenAlex English Titles & Abstracts (April 2025 Snapshot) This dataset is part of the PaperSeek framework, a semantic search engine designed for literature discovery using research questions and prior knowledge. PaperSeek is developed as part of a Master's thesis to explore novel approaches in enhancing academic search relevance. 📦 Dataset Overview Source: OpenAlex Snapshot Date: April 1st, 2025 Language: English Contents: Title Abstract Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Grozkal/PaperSeek-OpenAlex-Embeddings.textquestion-answering10M<n<100M5 likes632 downloads1y agoHugging Face06Voxel51 /fiftyone-embeddings-combined FiftyOne Embeddings Dataset This dataset combines the FiftyOne Q&A and function calling datasets with pre-computed embeddings for fast similarity search. Dataset Information Total samples: 28,118 Q&A samples: 14,069 Function samples: 14,049 Embedding model: text-embedding-3-large Embedding dimension: 3072 Schema query: The original question/query text response: The unified response content (either answer text for Q&A or function call text for function… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/fiftyone-embeddings-combined.textquestion-answering10K<n<100K1 likes416 downloads1y agoHugging Face07MongoDB /airbnb_embeddings Overview This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata. It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face. The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.tabularquestion-answering1K<n<10K7 likes401 downloads2y agoHugging Face08flax-sentence-embeddings /stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M20 likes343 downloads4y agoHugging Face09jspringer /open-synthetic-embeddingstextfeature-extraction1M<n<10M3 likes92 downloads1y agoHugging Face10calmgoose /book-embeddings Vector store of embeddings for books "1984" by George Orwell "The Almanac of Naval Ravikant" by Eric Jorgenson This is a faiss vector store created with instructor embeddings using LangChain . Use it for similarity search, question answering or anything else that leverages embeddings! 😃 Creating these embeddings can take a while so here's a convenient, downloadable one 🤗 How to use Specify the book from one of the following: "1984" "The Almanac of Naval Ravikant"… See the full description on the dataset page: https://huggingface.co/datasets/calmgoose/book-embeddings.question-answering11 likes68 downloads3y agoHugging Face11dhlak /legal_chroma_embeddings US & Iowa Legal Embeddings (RAG-Optimized) Dataset Summary A large-scale, high-quality embedding dataset covering U.S. federal law and Iowa state law, optimized for retrieval-augmented generation (RAG), semantic search, and multi-hop legal reasoning. The dataset contains ~3.13M embeddings generated with bge-m3, with chunking strategies specifically designed to preserve legal structure and maximize retrieval performance. Key highlights: ~3.13M embeddings across… See the full description on the dataset page: https://huggingface.co/datasets/dhlak/legal_chroma_embeddings.text-retrieval1M<n<10M0 likes62 downloads4mo agoHugging Face12fscheffczyk /20newsgroups_embeddings Dataset Card for feature vector embeddings of the 20newsgroup dataset Dataset Summary This dataset contains vector embeddings of the 20newsgroups dataset. The embeddings were created with the Sentence Transformers library using the multi-qa-MiniLM-L6-cos-v1 model. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data… See the full description on the dataset page: https://huggingface.co/datasets/fscheffczyk/20newsgroups_embeddings.tabularfeature-extraction10K<n<100K1 likes46 downloads4y agoHugging Face13nickmuchi /CFA_Level_1_Text_EmbeddingsVector store of embeddings for CFA Level 1 Curriculum This is a faiss vector store created with Sentence Transformer embeddings using LangChain . Use it for similarity search, question answering or anything else that leverages embeddings! 😃 Creating these embeddings can take a while so here's a convenient, downloadable one 🤗 How to use Download data Load to use with LangChain pip install -qqq langchain sentence_transformers faiss-cpu huggingface_hub import os from langchain.embeddings import… See the full description on the dataset page: https://huggingface.co/datasets/nickmuchi/CFA_Level_1_Text_Embeddings.question-answering3 likes44 downloads3y agoHugging Face14R0bk /open-australian-legal-embeddings-openaitext-retrieval1M<n<10M0 likes38 downloads2y agoHugging Face15cmpunkmannu /financebench-voyage-finance-2-embeddings FinanceBench (voyage-finance-2 embeddings) Pre-computed embeddings for the FinanceBench corpus. Skips ~$5-15 of Voyage API cost and ~30 minutes of ingest time vs re-embedding from raw PDFs. Intended consumer: the RAG agent at Rishabhmannu/financebench-rag-agent (install: pip install financebench-rag-agent). What's in the box (frozen) Field Value Source corpus FinanceBench (SEC filings: 10-K, 10-Q, 8-K, earnings releases) Parser pypdf (canonical)… See the full description on the dataset page: https://huggingface.co/datasets/cmpunkmannu/financebench-voyage-finance-2-embeddings.tabularsentence-similarity10K<n<100K1 likes35 downloads4mo agoHugging Face16fscheffczyk /2D_20newsgroups_embeddings Dataset Card for feature vector embeddings of the 20newsgroup dataset Dataset Summary This dataset contains dimensional reduced vector embeddings of the 20newsgroups dataset. This dataset contains two dimensions. The dimensional reduced embeddings were created with the TruncatedSVD function from the scikit-learn library. These reduced feature vectors are based on the fscheffczyk/20newsgroup_embeddings dataset. Supported Tasks and Leaderboards [More… See the full description on the dataset page: https://huggingface.co/datasets/fscheffczyk/2D_20newsgroups_embeddings.tabularfeature-extraction10K<n<100K1 likes32 downloads4y agoHugging Face17biznetgio /indonesia-law-qa-embeddingstextquestion-answering1K<n<10K6 likes31 downloads2y agoHugging Face18BassemE /mulesoft-documentation-embeddings mulesoft-documentation-embeddings MuleSoft Documentation Embeddings for RAG Applications Dataset Information Version: 1.0.0 Created: 2025-09-16T02:41:16.352809 Source: Vector Database License: MIT Language: en Task Categories question-answering, retrieval, knowledge-base Dataset Statistics SkillPilotDataSet_v11 Total Objects: 6430 Unique Properties: 13 Knowledge Sources: mulesoft, user_defined_docs Average Content Length: 5079… See the full description on the dataset page: https://huggingface.co/datasets/BassemE/mulesoft-documentation-embeddings.tabularquestion-answering1K<n<10K0 likes29 downloads1y agoHugging Face19Mercity /ramayana-embeddingsRamayana Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Ramayana, suitable for various NLP tasks, semantic search, and text retrieval purposes. Dataset relesased by Mercity AI! Dataset SourceThe original textual data has been sourced from the repository: Sanskrit Sahitya Data Repository We extend our sincere gratitude to the maintainers of this repository for compiling and sharing valuable Sanskrit literature datasets openly. Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/ramayana-embeddings.textquestion-answering10K<n<100K1 likes28 downloads1y agoHugging Face20Mercity /mahabharat-embeddingsMahabharat Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Mahabharat, suitable for various NLP tasks, semantic search, and text retrieval purposes. Dataset relesased by Mercity AI! Dataset SourceThe original textual data has been sourced from the repository: Sanskrit Sahitya Data Repository We extend our sincere gratitude to the maintainers of this repository for compiling and sharing valuable Sanskrit literature datasets openly.… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/mahabharat-embeddings.textquestion-answering10K<n<100K1 likes27 downloads1y agoHugging Face21HFforLegal /embedding-models Reference models for integration into HF for Legal 🤗 This dataset comprises a collection of models aimed at streamlining and partially automating the embedding process. Each model entry within this dataset includes essential information such as model identifiers, embedding configurations, and specific parameters, ensuring that users can seamlessly integrate these models into their workflows with minimal setup and maximum efficiency. Dataset Structure Field Type… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/embedding-models.tabulartabular-to-textn<1K3 likes26 downloads2y agoHugging Face22svetfm /epstein-files-nov11-25-house-post-ocr-embeddings Epstein Files Document Embeddings Text embeddings generated from the House Oversight Committee's Epstein document release. Source Dataset This dataset is derived from: tensonaut/EPSTEIN_FILES_20K The source dataset contains OCR'd text from the original House Oversight Committee PDF release. Dataset Structure Field Type Description source_file string Source document filename chunk_index int Position of chunk within document text string Original… See the full description on the dataset page: https://huggingface.co/datasets/svetfm/epstein-files-nov11-25-house-post-ocr-embeddings.texttext-retrieval10K<n<100K3 likes26 downloads10mo agoHugging Face23Laz4rz /wikipedia_stem_small_rag_embeddings STEMWikiSmallRAG with embeddings This dataset contains wikipedia entries from STEM field, unfortunately there is also Business&Economics... but I thought it may contain some useful data as well, even by accident. Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_stem_small_rag_embeddings.texttext-generation100K<n<1M0 likes22 downloads2y agoHugging Face24davenporten /nrc-regulatory-embeddings NRC Regulatory Embeddings 37,734 chunked and embedded NRC nuclear regulatory documents, ready for use in RAG pipelines. Built for the nrc-licensing-rag project, an AI system for analyzing nuclear Combined License Applications (COLAs). Contents Source Documents NUREG-0800 (Standard Review Plan) chapters 1-19 2,436 sections 10 CFR Parts 20, 50, 51, 52, 72, 73, 100 ~504 sections Regulatory Guide Division 1 (1.1-1.262) 242 guides Regulatory Guide Division 4… See the full description on the dataset page: https://huggingface.co/datasets/davenporten/nrc-regulatory-embeddings.texttext-retrieval10K<n<100K2 likes22 downloads6mo agoHugging Face25kevinyulistian /indonesia-law-qa-embeddingstextquestion-answering1K<n<10K0 likes22 downloads4mo agoHugging Face26andfrca /medpt-qwen3-embeddings MedPT with Qwen3 Embeddings About This Derived Dataset This repository is a derived version of the original MedPT dataset, enriched with precomputed semantic embeddings generated from the question field using Qwen3-Embedding-0.6B. The original MedPT corpus, including its medical questions, answers, metadata, collection methodology, curation procedures, annotations, benchmarks, and scientific contributions, was created by Farber et al. (2026). This repository does… See the full description on the dataset page: https://huggingface.co/datasets/andfrca/medpt-qwen3-embeddings.textquestion-answering100K<n<1M0 likes17 downloads1mo agoHugging Face27Singhchandann /finanical-rag-embedding-dataset_marathigated Finanical-Rag-Embedding-Dataset Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The Finanical-Rag-Embedding-Dataset Marathi dataset is a meticulously curated collection of 6998 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/finanical-rag-embedding-dataset_marathi.texttext-classification1K<n<10K0 likes12 downloads1y agoHugging Face28hi-zero /pubmed_QA_embeddingThis embedded data and original data are came from (https://huggingface.co/datasets/qiaojin/PubMedQA), pqa_artificial subset used for PubMed QA test set. You can get original Pubmed QA data by following above link. "embeddings" columns are made by following code lines from sentence_transformers import SentenceTransformer ST = SentenceTransformer("mixedbread-ai/mxbai-embed-large-v1") def data_preprocess(examples) : context_dic = examples['context'] total_con = '' for i in… See the full description on the dataset page: https://huggingface.co/datasets/hi-zero/pubmed_QA_embedding.texttext-generation1K<n<10K0 likes11 downloads2y agoHugging Face29likhitjuttada /ft-embeddingmodel-RAG-dataset Dataset Card for Dataset Name This dataset aims to be a base template for fine-tuning embedding models for enhanced retrieval performance in RAG pipelines. It has been generated locally using Mistral:7B on Ollama using a simple prompt that prompts the model to generate 5 questions for each document chunk of Apple's Environmental Progress Report 2024 Dataset Details Dataset Description Curated by: Likhit Juttada Funded by [optional]: NA Credits… See the full description on the dataset page: https://huggingface.co/datasets/likhitjuttada/ft-embeddingmodel-RAG-dataset.textquestion-answeringn<1K0 likes10 downloads11mo agoHugging Face30jaiw /lex_fridman_podcast_embeddings Description This dataset contains csv's from the Lex Fridman podcast transcripts provided by Whispering-GPT. I split the episode transcripts into parent and child chunks for use with RAG. The parent chunks is size 500 and the child is size 50. The children come with embeddings using OpenAI text-embedding-3-small with 1024 dimensionality. Motivation This was designed for use with ParentDocumentRetriever or LlamaIndex. It should provide better retrievals for queries on… See the full description on the dataset page: https://huggingface.co/datasets/jaiw/lex_fridman_podcast_embeddings.question-answering0 likes9 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.