datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineWeb2-embedded
FineWeb2-embedded
Dataset summary
FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research.
Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer texts… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.wdc-common-crawl-embedded-jsonldopenwebtext-t5refinedweb-embedded_prototypeContaining 768-dimensional embedding vectors derived from the "content" column of the RefinedWeb dataset.
The embeddings were generated using the E5-Base-4k model with a context length of 1024, employing scaled dot-product attention.
Original dataset:
This dataset bases as derivative work on RefinedWeb, an English web dataset created by the TII (Technology Innovation Institute).
Attribution is given to the TII as original authors of the RefinedWeb dataset as per Section 4 of RefinedWeb's Open… See the full description on the dataset page: https://huggingface.co/datasets/Marcus2112/refinedweb-embedded_prototype.xsum_validation_t5Wikipedia-TR-2023-Embedded-Dump
Wikipedia-TR-2023-Embedded-Dump
Türkçe Vikipedi (tr.wikipedia.org) makaleleri, retrieval / RAG kullanım
senaryoları için parçalanmış (chunk) ve embedding'lenmiş hâliyle. Her
makale bir parent (ana) chunk'a (tüm makale metni, bağlam
genişletmek için) ve birden fazla child (alt) chunk'a (her biri
kendi embedding'ine sahip küçük pasajlar) bölünmüştür.
İçerik
Makale
348.751
Embedding'li child chunk
1.308.623
Parent chunk (embeddingsiz)
348.751… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Wikipedia-TR-2023-Embedded-Dump.embedded-topology
Exp 3B: Embedded topology in trained recurrent operators
A 14,553-configuration computational sweep of recurrent network
architectures trained on dynamical systems. Each configuration's
hidden-state activations were measured for topological fidelity to
the driving system using persistent homology and Gauss linking
integrals. Six post-hoc analyses on saved checkpoints probe the
operator properties the embedding theorems describe abstractly.
This dataset is the empirical companion… See the full description on the dataset page: https://huggingface.co/datasets/heights1976/embedded-topology.CT-RATE_RAPTOR_DINOV3_Embedded_Validwmt14_de-en_validation_t5embedded_textxsum_train_t5UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6bdetails_EmbeddedLLM__Mistral-7B-Merge-14-v0.2
Dataset Card for Evaluation run of EmbeddedLLM/Mistral-7B-Merge-14-v0.2
Dataset automatically created during the evaluation run of model EmbeddedLLM/Mistral-7B-Merge-14-v0.2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EmbeddedLLM__Mistral-7B-Merge-14-v0.2.paired-open-images-embedded-pe-core-g14-448
Paired Open Images with PE-Core-G14-448 Embeddings
This dataset contains pairs of images from Open Images along with their embeddings computed using Meta's Perception Encoder (PE-Core-G14-448).
Each row contains two images (as JPEG bytes), their metadata, and their corresponding 1280-dimensional embeddings.
Data Layout
Column
Description
image1_jpeg
JPEG bytes for the first image
image1_metadata
Metadata for the first image
image2_jpeg
JPEG bytes for… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-open-images-embedded-pe-core-g14-448.embedded_movies
sample_mflix.embedded_movies
This data set contains details on movies with genres of Western, Action, or Fantasy. Each document contains a single movie, and information such as its title, release year, and cast.
In addition, documents in this collection include a plot_embedding field that contains embeddings created using OpenAI's text-embedding-ada-002 embedding model that you can use with the Atlas Search vector search feature.
Overview
This dataset offers a… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/embedded_movies.UltraData-Math-L2-preview-embedded-with-idsUltraData-Math-L3-Textbook-Exercise-Synthetic-split-qwen3-0.6b-embeddedchatdoctor-embedded
Chat Doctor with Embeddings
This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped:
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
414k
Token Count
1.7b
Origin
https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view
Source of raw data
?
Processing details
paper
Embedding Model
BAAI/bge-small-en-v1.5
Data Diversity
index
Example Output
GPT-4 Rationale
GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.LLaVA-CC3M-Pretrain-595K-Embedded
Dataset derived from liuhaotian/LLaVA-CC3M-Pretrain-595K
Dataset details
Dataset type:
LLaVA Visual Instruct CC3M Pretrain 595K is a subset of CC-3M dataset, filtered with a more balanced concept coverage distribution.
Captions are also associated with BLIP synthetic caption for reference.
It is constructed for the pretraining stage for feature alignment in visual instruction tuning.
We aim to build large multimodal towards GPT-4 vision/language capability.
20newsgroups_embedded
Dataset Card for 20-Newsgroups Embedded
This provides a subset of 20-Newsgroup posts, along with sentence embeddings, and a dimension reduced 2D data map.
This provides a basic setup for experimentation with various neural topic modelling approaches.
Dataset Details
Dataset Description
This is a dataset containing posts from the classic 20-Newsgroups dataset, along with sentence embeddings, and a dimension reduced 2D data map.
Per the source:
The… See the full description on the dataset page: https://huggingface.co/datasets/lmcinnes/20newsgroups_embedded.SlimPajama-6B-embedded
Dataset Card for SlimPajama-6B-embedded
This is a copy of DKYoon/SlimPajama-6B, together with embeddings generated by thenlper/gte-large.
There are 5.49 million examples of text, a representative random sample of SlimPajama-627B. Each text is associated with a 1024-dimensional embedding vector that is meant to represent the semantic content. The vectors were generated by average-pooling (max-pooling dataset to come in the future).
This dataset is intended to help with downstream… See the full description on the dataset page: https://huggingface.co/datasets/sproos/SlimPajama-6B-embedded.wmt14_de-en_train_t5miriad-embeddedforlinx-downloads
Forlinx Embedded: AI Models & Multimedia Resource Hub
Welcome to the official Forlinx Embedded resource repository. This hub serves as a centralized ecosystem providing high-performance Edge AI models, Hardware Development Kits, and Creative Marketing Assets specifically tailored for our SoM (System on Module) and SBC (Single Board Computer) platforms.
Resource Categories
To accelerate your deployment and content creation, our resources are organized into three… See the full description on the dataset page: https://huggingface.co/datasets/forlinx-embedded/forlinx-downloads.UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6b_newembedded_1
Dataset Card for "embedded_1"
More Information needed
details_EmbeddedLLM__Mistral-7B-Merge-14-v0
Dataset Card for Evaluation run of EmbeddedLLM/Mistral-7B-Merge-14-v0
Dataset automatically created during the evaluation run of model EmbeddedLLM/Mistral-7B-Merge-14-v0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EmbeddedLLM__Mistral-7B-Merge-14-v0.airep-embedded-evaluation-profile
AIREP Embedded Evaluation Profile v0.1
This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an
experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical
specification history lives in the AIREP GitHub repository. Byte identity between this mirror and
its source commit is a distribution-integrity property, not independent scientific verification.
Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.embedded_dataset_domain_similarity_EleutherAI_gpt-neo-125mlamini-embedded-instructions-only
Dataset Card for "lamini-embedded-instructions-only"
More Information needed
