datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineWeb2-embedded
FineWeb2-embedded
Dataset summary
FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research.
Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer texts… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.wdc-common-crawl-embedded-jsonldrefinedweb-embedded_prototypeContaining 768-dimensional embedding vectors derived from the "content" column of the RefinedWeb dataset.
The embeddings were generated using the E5-Base-4k model with a context length of 1024, employing scaled dot-product attention.
Original dataset:
This dataset bases as derivative work on RefinedWeb, an English web dataset created by the TII (Technology Innovation Institute).
Attribution is given to the TII as original authors of the RefinedWeb dataset as per Section 4 of RefinedWeb's Open… See the full description on the dataset page: https://huggingface.co/datasets/Marcus2112/refinedweb-embedded_prototype.Wikipedia-TR-2023-Embedded-Dump
Wikipedia-TR-2023-Embedded-Dump
Türkçe Vikipedi (tr.wikipedia.org) makaleleri, retrieval / RAG kullanım
senaryoları için parçalanmış (chunk) ve embedding'lenmiş hâliyle. Her
makale bir parent (ana) chunk'a (tüm makale metni, bağlam
genişletmek için) ve birden fazla child (alt) chunk'a (her biri
kendi embedding'ine sahip küçük pasajlar) bölünmüştür.
İçerik
Makale
348.751
Embedding'li child chunk
1.308.623
Parent chunk (embeddingsiz)
348.751… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Wikipedia-TR-2023-Embedded-Dump.CT-RATE_RAPTOR_DINOV3_Embedded_Validxsum_train_t5embedded_textUltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6bUltraData-Math-L2-preview-embedded-with-idsUltraData-Math-L3-Textbook-Exercise-Synthetic-split-qwen3-0.6b-embeddedchatdoctor-embedded
Chat Doctor with Embeddings
This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped:
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
414k
Token Count
1.7b
Origin
https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view
Source of raw data
?
Processing details
paper
Embedding Model
BAAI/bge-small-en-v1.5
Data Diversity
index
Example Output
GPT-4 Rationale
GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.embedded_movies
sample_mflix.embedded_movies
This data set contains details on movies with genres of Western, Action, or Fantasy. Each document contains a single movie, and information such as its title, release year, and cast.
In addition, documents in this collection include a plot_embedding field that contains embeddings created using OpenAI's text-embedding-ada-002 embedding model that you can use with the Atlas Search vector search feature.
Overview
This dataset offers a… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/embedded_movies.SlimPajama-6B-embedded
Dataset Card for SlimPajama-6B-embedded
This is a copy of DKYoon/SlimPajama-6B, together with embeddings generated by thenlper/gte-large.
There are 5.49 million examples of text, a representative random sample of SlimPajama-627B. Each text is associated with a 1024-dimensional embedding vector that is meant to represent the semantic content. The vectors were generated by average-pooling (max-pooling dataset to come in the future).
This dataset is intended to help with downstream… See the full description on the dataset page: https://huggingface.co/datasets/sproos/SlimPajama-6B-embedded.20newsgroups_embedded
Dataset Card for 20-Newsgroups Embedded
This provides a subset of 20-Newsgroup posts, along with sentence embeddings, and a dimension reduced 2D data map.
This provides a basic setup for experimentation with various neural topic modelling approaches.
Dataset Details
Dataset Description
This is a dataset containing posts from the classic 20-Newsgroups dataset, along with sentence embeddings, and a dimension reduced 2D data map.
Per the source:
The… See the full description on the dataset page: https://huggingface.co/datasets/lmcinnes/20newsgroups_embedded.miriad-embeddedUltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6b_newembedded_1
Dataset Card for "embedded_1"
More Information needed
airep-embedded-evaluation-profile
AIREP Embedded Evaluation Profile v0.1
This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an
experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical
specification history lives in the AIREP GitHub repository. Byte identity between this mirror and
its source commit is a distribution-integrity property, not independent scientific verification.
Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.lamini-embedded-instructions-only
Dataset Card for "lamini-embedded-instructions-only"
More Information needed
corpus_1_embedded_deduplicated
Dataset Card for "corpus_1_embedded_deduplicated"
More Information needed
lamini-embedded
Dataset Card for "lamini-embedded"
More Information needed
synthetic-clinical-notes-embedded
Synthetic Clinical Notes
This dataset is post-processed version of starmpcc/Asclepius-Synthetic-Clinical-Notes:
Turn into Alpaca format (instruction, input, and output)
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
158k
Token Count
648m
Origin
https://figshare.com/authors/Zhengyun_Zhao/16480335
Source of raw data
PubMed Central (PMC) and MIMIC 3
Processing details
original, paper
Embedding Model… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/synthetic-clinical-notes-embedded.embedded_datasets_0822
Dataset Card for "combined_embedded_v2"
More Information needed
github-embedded
github-embedded
A bunch of python code, gotten from Github, along with a json dump of their ast's and a vector embedding of the code using openai's text-embedding-3-small.
donaroma3421.42_embeddedkill-life-embedded-qa
Ailiance — Kill-LIFE Embedded Knowledge Base
🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/kill-life-embedded-qa. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025).
Knowledge-base Q&A spécifique au projet Kill_LIFE (compagnon vocal embarqué ESP32-S3 + Mascarade) : composants matériels, schémas KiCad du board minimal, simulations SPICE de l'alimentation/I2C/I2S/audio, et architecture du firmware C++ (pipeline… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/kill-life-embedded-qa.embeddedflanorca_embeddedembedded-systems-qa
Embedded Systems Engineering Q&A — Instruction Dataset
A hand-authored instruction-tuning dataset of technical question/answer pairs for
embedded systems engineering, formatted for supervised fine-tuning of Mistral 7B
(Alpaca-style instruction / input / output schema).
At a glance
Entries
302
Format
JSONL, one JSON object per line
Schema
{"instruction": <question>, "input": "", "output": <answer>}
Language
English
Avg. answer length
~590… See the full description on the dataset page: https://huggingface.co/datasets/eniomecaj/embedded-systems-qa.BlueLionEcho_tran_embedded_filtered
