datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tech-news-embeddings
Overview
HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023.
To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256.
Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.PaperSeek-OpenAlex-Embeddings
📚 PaperSeek: OpenAlex English Titles & Abstracts (April 2025 Snapshot)
This dataset is part of the PaperSeek framework, a semantic search engine designed for literature discovery using research questions and prior knowledge. PaperSeek is developed as part of a Master's thesis to explore novel approaches in enhancing academic search relevance.
📦 Dataset Overview
Source: OpenAlex
Snapshot Date: April 1st, 2025
Language: English
Contents:
Title
Abstract
Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Grozkal/PaperSeek-OpenAlex-Embeddings.fiftyone-embeddings-combined
FiftyOne Embeddings Dataset
This dataset combines the FiftyOne Q&A and function calling datasets with pre-computed embeddings for fast similarity search.
Dataset Information
Total samples: 28,118
Q&A samples: 14,069
Function samples: 14,049
Embedding model: text-embedding-3-large
Embedding dimension: 3072
Schema
query: The original question/query text
response: The unified response content (either answer text for Q&A or function call text for function… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/fiftyone-embeddings-combined.airbnb_embeddings
Overview
This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata.
It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face.
The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.open-synthetic-embeddingsbook-embeddings
Vector store of embeddings for books
"1984" by George Orwell
"The Almanac of Naval Ravikant" by Eric Jorgenson
This is a faiss vector store created with instructor embeddings using LangChain . Use it for similarity search, question answering or anything else that leverages embeddings! 😃
Creating these embeddings can take a while so here's a convenient, downloadable one 🤗
How to use
Specify the book from one of the following:
"1984"
"The Almanac of Naval Ravikant"… See the full description on the dataset page: https://huggingface.co/datasets/calmgoose/book-embeddings.legal_chroma_embeddings
US & Iowa Legal Embeddings (RAG-Optimized)
Dataset Summary
A large-scale, high-quality embedding dataset covering U.S. federal law and Iowa state law, optimized for retrieval-augmented generation (RAG), semantic search, and multi-hop legal reasoning.
The dataset contains ~3.13M embeddings generated with bge-m3, with chunking strategies specifically designed to preserve legal structure and maximize retrieval performance.
Key highlights:
~3.13M embeddings across… See the full description on the dataset page: https://huggingface.co/datasets/dhlak/legal_chroma_embeddings.20newsgroups_embeddings
Dataset Card for feature vector embeddings of the 20newsgroup dataset
Dataset Summary
This dataset contains vector embeddings of the 20newsgroups dataset.
The embeddings were created with the Sentence Transformers library using the multi-qa-MiniLM-L6-cos-v1 model.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data… See the full description on the dataset page: https://huggingface.co/datasets/fscheffczyk/20newsgroups_embeddings.CFA_Level_1_Text_EmbeddingsVector store of embeddings for CFA Level 1 Curriculum
This is a faiss vector store created with Sentence Transformer embeddings using LangChain . Use it for similarity search, question answering or anything else that leverages embeddings! 😃
Creating these embeddings can take a while so here's a convenient, downloadable one 🤗
How to use
Download data
Load to use with LangChain
pip install -qqq langchain sentence_transformers faiss-cpu huggingface_hub
import os
from langchain.embeddings import… See the full description on the dataset page: https://huggingface.co/datasets/nickmuchi/CFA_Level_1_Text_Embeddings.open-australian-legal-embeddings-openaifinancebench-voyage-finance-2-embeddings
FinanceBench (voyage-finance-2 embeddings)
Pre-computed embeddings for the FinanceBench corpus. Skips ~$5-15 of Voyage API cost and ~30 minutes of ingest time vs re-embedding from raw PDFs. Intended consumer: the RAG agent at Rishabhmannu/financebench-rag-agent (install: pip install financebench-rag-agent).
What's in the box (frozen)
Field
Value
Source corpus
FinanceBench (SEC filings: 10-K, 10-Q, 8-K, earnings releases)
Parser
pypdf (canonical)… See the full description on the dataset page: https://huggingface.co/datasets/cmpunkmannu/financebench-voyage-finance-2-embeddings.2D_20newsgroups_embeddings
Dataset Card for feature vector embeddings of the 20newsgroup dataset
Dataset Summary
This dataset contains dimensional reduced vector embeddings of the 20newsgroups dataset. This dataset contains two dimensions.
The dimensional reduced embeddings were created with the TruncatedSVD function from the scikit-learn library.
These reduced feature vectors are based on the fscheffczyk/20newsgroup_embeddings dataset.
Supported Tasks and Leaderboards
[More… See the full description on the dataset page: https://huggingface.co/datasets/fscheffczyk/2D_20newsgroups_embeddings.indonesia-law-qa-embeddingsmulesoft-documentation-embeddings
mulesoft-documentation-embeddings
MuleSoft Documentation Embeddings for RAG Applications
Dataset Information
Version: 1.0.0
Created: 2025-09-16T02:41:16.352809
Source: Vector Database
License: MIT
Language: en
Task Categories
question-answering, retrieval, knowledge-base
Dataset Statistics
SkillPilotDataSet_v11
Total Objects: 6430
Unique Properties: 13
Knowledge Sources: mulesoft, user_defined_docs
Average Content Length: 5079… See the full description on the dataset page: https://huggingface.co/datasets/BassemE/mulesoft-documentation-embeddings.ramayana-embeddingsRamayana Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Ramayana, suitable for various NLP tasks, semantic search, and text retrieval purposes.
Dataset relesased by Mercity AI!
Dataset SourceThe original textual data has been sourced from the repository: Sanskrit Sahitya Data Repository
We extend our sincere gratitude to the maintainers of this repository for compiling and sharing valuable Sanskrit literature datasets openly.
Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/ramayana-embeddings.mahabharat-embeddingsMahabharat Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Mahabharat, suitable for various NLP tasks, semantic search, and text retrieval purposes.
Dataset relesased by Mercity AI!
Dataset SourceThe original textual data has been sourced from the repository: Sanskrit Sahitya Data Repository
We extend our sincere gratitude to the maintainers of this repository for compiling and sharing valuable Sanskrit literature datasets openly.… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/mahabharat-embeddings.embedding-models
Reference models for integration into HF for Legal 🤗
This dataset comprises a collection of models aimed at streamlining and partially automating the embedding process. Each model entry within this dataset includes essential information such as model identifiers, embedding configurations, and specific parameters, ensuring that users can seamlessly integrate these models into their workflows with minimal setup and maximum efficiency.
Dataset Structure
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/embedding-models.epstein-files-nov11-25-house-post-ocr-embeddings
Epstein Files Document Embeddings
Text embeddings generated from the House Oversight Committee's Epstein document release.
Source Dataset
This dataset is derived from: tensonaut/EPSTEIN_FILES_20K
The source dataset contains OCR'd text from the original House Oversight Committee PDF release.
Dataset Structure
Field
Type
Description
source_file
string
Source document filename
chunk_index
int
Position of chunk within document
text
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/svetfm/epstein-files-nov11-25-house-post-ocr-embeddings.wikipedia_stem_small_rag_embeddings
STEMWikiSmallRAG with embeddings
This dataset contains wikipedia entries from STEM field, unfortunately there is also Business&Economics... but I thought it may contain some useful data as well, even by accident.
Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_stem_small_rag_embeddings.nrc-regulatory-embeddings
NRC Regulatory Embeddings
37,734 chunked and embedded NRC nuclear regulatory documents, ready for use in RAG pipelines.
Built for the nrc-licensing-rag project, an AI system for analyzing nuclear Combined License Applications (COLAs).
Contents
Source
Documents
NUREG-0800 (Standard Review Plan) chapters 1-19
2,436 sections
10 CFR Parts 20, 50, 51, 52, 72, 73, 100
~504 sections
Regulatory Guide Division 1 (1.1-1.262)
242 guides
Regulatory Guide Division 4… See the full description on the dataset page: https://huggingface.co/datasets/davenporten/nrc-regulatory-embeddings.indonesia-law-qa-embeddingsmedpt-qwen3-embeddings
MedPT with Qwen3 Embeddings
About This Derived Dataset
This repository is a derived version of the original MedPT dataset, enriched with precomputed semantic embeddings generated from the question field using Qwen3-Embedding-0.6B.
The original MedPT corpus, including its medical questions, answers, metadata, collection methodology, curation procedures, annotations, benchmarks, and scientific contributions, was created by Farber et al. (2026).
This repository does… See the full description on the dataset page: https://huggingface.co/datasets/andfrca/medpt-qwen3-embeddings.finanical-rag-embedding-dataset_marathi
Finanical-Rag-Embedding-Dataset Marathi Dataset: High-Quality Marathi NLP Corpus
📌 Overview
The Finanical-Rag-Embedding-Dataset Marathi dataset is a meticulously curated collection of 6998 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness.
This dataset is designed for semantic search, text classification, and various NLP tasks, making it a… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/finanical-rag-embedding-dataset_marathi.pubmed_QA_embeddingThis embedded data and original data are came from (https://huggingface.co/datasets/qiaojin/PubMedQA), pqa_artificial subset used for PubMed QA test set.
You can get original Pubmed QA data by following above link.
"embeddings" columns are made by following code lines
from sentence_transformers import SentenceTransformer
ST = SentenceTransformer("mixedbread-ai/mxbai-embed-large-v1")
def data_preprocess(examples) :
context_dic = examples['context']
total_con = ''
for i in… See the full description on the dataset page: https://huggingface.co/datasets/hi-zero/pubmed_QA_embedding.ft-embeddingmodel-RAG-dataset
Dataset Card for Dataset Name
This dataset aims to be a base template for fine-tuning embedding models for enhanced retrieval performance in RAG pipelines.
It has been generated locally using Mistral:7B on Ollama using a simple prompt that prompts the model to generate 5 questions for each document chunk of Apple's Environmental Progress Report 2024
Dataset Details
Dataset Description
Curated by: Likhit Juttada
Funded by [optional]: NA
Credits… See the full description on the dataset page: https://huggingface.co/datasets/likhitjuttada/ft-embeddingmodel-RAG-dataset.lex_fridman_podcast_embeddings
Description
This dataset contains csv's from the Lex Fridman podcast transcripts provided by Whispering-GPT.
I split the episode transcripts into parent and child chunks for use with RAG.
The parent chunks is size 500 and the child is size 50. The children come with embeddings using OpenAI text-embedding-3-small with 1024 dimensionality.
Motivation
This was designed for use with ParentDocumentRetriever or LlamaIndex.
It should provide better retrievals for queries on… See the full description on the dataset page: https://huggingface.co/datasets/jaiw/lex_fridman_podcast_embeddings.
