datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tech-news-embeddings
Overview
HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023.
To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256.
Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.PaperSeek-OpenAlex-Embeddings
📚 PaperSeek: OpenAlex English Titles & Abstracts (April 2025 Snapshot)
This dataset is part of the PaperSeek framework, a semantic search engine designed for literature discovery using research questions and prior knowledge. PaperSeek is developed as part of a Master's thesis to explore novel approaches in enhancing academic search relevance.
📦 Dataset Overview
Source: OpenAlex
Snapshot Date: April 1st, 2025
Language: English
Contents:
Title
Abstract
Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Grozkal/PaperSeek-OpenAlex-Embeddings.fiftyone-embeddings-combined
FiftyOne Embeddings Dataset
This dataset combines the FiftyOne Q&A and function calling datasets with pre-computed embeddings for fast similarity search.
Dataset Information
Total samples: 28,118
Q&A samples: 14,069
Function samples: 14,049
Embedding model: text-embedding-3-large
Embedding dimension: 3072
Schema
query: The original question/query text
response: The unified response content (either answer text for Q&A or function call text for function… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/fiftyone-embeddings-combined.airbnb_embeddings
Overview
This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata.
It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face.
The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.open-synthetic-embeddingsfinancebench-voyage-finance-2-embeddings
FinanceBench (voyage-finance-2 embeddings)
Pre-computed embeddings for the FinanceBench corpus. Skips ~$5-15 of Voyage API cost and ~30 minutes of ingest time vs re-embedding from raw PDFs. Intended consumer: the RAG agent at Rishabhmannu/financebench-rag-agent (install: pip install financebench-rag-agent).
What's in the box (frozen)
Field
Value
Source corpus
FinanceBench (SEC filings: 10-K, 10-Q, 8-K, earnings releases)
Parser
pypdf (canonical)… See the full description on the dataset page: https://huggingface.co/datasets/cmpunkmannu/financebench-voyage-finance-2-embeddings.indonesia-law-qa-embeddingsramayana-embeddingsRamayana Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Ramayana, suitable for various NLP tasks, semantic search, and text retrieval purposes.
Dataset relesased by Mercity AI!
Dataset SourceThe original textual data has been sourced from the repository: Sanskrit Sahitya Data Repository
We extend our sincere gratitude to the maintainers of this repository for compiling and sharing valuable Sanskrit literature datasets openly.
Embedding… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/ramayana-embeddings.mulesoft-documentation-embeddings
mulesoft-documentation-embeddings
MuleSoft Documentation Embeddings for RAG Applications
Dataset Information
Version: 1.0.0
Created: 2025-09-16T02:41:16.352809
Source: Vector Database
License: MIT
Language: en
Task Categories
question-answering, retrieval, knowledge-base
Dataset Statistics
SkillPilotDataSet_v11
Total Objects: 6430
Unique Properties: 13
Knowledge Sources: mulesoft, user_defined_docs
Average Content Length: 5079… See the full description on the dataset page: https://huggingface.co/datasets/BassemE/mulesoft-documentation-embeddings.epstein-files-nov11-25-house-post-ocr-embeddings
Epstein Files Document Embeddings
Text embeddings generated from the House Oversight Committee's Epstein document release.
Source Dataset
This dataset is derived from: tensonaut/EPSTEIN_FILES_20K
The source dataset contains OCR'd text from the original House Oversight Committee PDF release.
Dataset Structure
Field
Type
Description
source_file
string
Source document filename
chunk_index
int
Position of chunk within document
text
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/svetfm/epstein-files-nov11-25-house-post-ocr-embeddings.nrc-regulatory-embeddings
NRC Regulatory Embeddings
37,734 chunked and embedded NRC nuclear regulatory documents, ready for use in RAG pipelines.
Built for the nrc-licensing-rag project, an AI system for analyzing nuclear Combined License Applications (COLAs).
Contents
Source
Documents
NUREG-0800 (Standard Review Plan) chapters 1-19
2,436 sections
10 CFR Parts 20, 50, 51, 52, 72, 73, 100
~504 sections
Regulatory Guide Division 1 (1.1-1.262)
242 guides
Regulatory Guide Division 4… See the full description on the dataset page: https://huggingface.co/datasets/davenporten/nrc-regulatory-embeddings.wikipedia_stem_small_rag_embeddings
STEMWikiSmallRAG with embeddings
This dataset contains wikipedia entries from STEM field, unfortunately there is also Business&Economics... but I thought it may contain some useful data as well, even by accident.
Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_stem_small_rag_embeddings.mahabharat-embeddingsMahabharat Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Mahabharat, suitable for various NLP tasks, semantic search, and text retrieval purposes.
Dataset relesased by Mercity AI!
Dataset SourceThe original textual data has been sourced from the repository: Sanskrit Sahitya Data Repository
We extend our sincere gratitude to the maintainers of this repository for compiling and sharing valuable Sanskrit literature datasets openly.… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/mahabharat-embeddings.indonesia-law-qa-embeddingsmedpt-qwen3-embeddings
MedPT with Qwen3 Embeddings
About This Derived Dataset
This repository is a derived version of the original MedPT dataset, enriched with precomputed semantic embeddings generated from the question field using Qwen3-Embedding-0.6B.
The original MedPT corpus, including its medical questions, answers, metadata, collection methodology, curation procedures, annotations, benchmarks, and scientific contributions, was created by Farber et al. (2026).
This repository does… See the full description on the dataset page: https://huggingface.co/datasets/andfrca/medpt-qwen3-embeddings.epstein-emails-embeddingsdictionary-embeddingsEmbeddings generated from the model multi-qa-mpnet-base-dot-v1 being trained on MAKILINGDING/english_dictionary
bhagavad_gita-embeddingsBhagavad Gita Embedding DatasetThis repository contains a specialized embedding dataset of the ancient Indian epic, Bhagavad Gita, suitable for various NLP tasks, semantic search, and text retrieval purposes.
Dataset relesased by Mercity AI!
Dataset SourceThe original textual data has been sourced from the repository: VedaBase.io - Bhagavad Gita Library by ISKCON
We extend our sincere gratitude to ISKCON and the maintainers of VedaBase.io for compiling and openly sharing valuable Sanskrit… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/bhagavad_gita-embeddings.epstein-files-nov11-25-house-post-ocr-embeddings
Epstein Files Document Embeddings
Text embeddings generated from the House Oversight Committee's Epstein document release.
Source Dataset
This dataset is derived from: tensonaut/EPSTEIN_FILES_20K
The source dataset contains OCR'd text from the original House Oversight Committee PDF release.
Dataset Structure
Field
Type
Description
source_file
string
Source document filename
chunk_index
int
Position of chunk within document
text
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/TinaWild09/epstein-files-nov11-25-house-post-ocr-embeddings.airbnb_embeddings
Overview
This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata.
It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face.
The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/Maki-99/airbnb_embeddings.
