datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ALPHAGenome-Embeddings
ALPHAGenome hg38 Embeddings
Pre-computed ALPHAGenome DNA foundation model embeddings for the entire human genome (hg38 / GRCh38).
The human genome is divided into ~22,000 non-overlapping 131 KB bins. Each bin's DNA sequence is embedded into a 3,072-dimensional latent space using the ALPHAGenome foundation model.
Companion project
These embeddings power the ALPHAGenome UMAP Explorer — an interactive browser visualization of latent relationships between genomic regions:
→… See the full description on the dataset page: https://huggingface.co/datasets/lagosproject/ALPHAGenome-Embeddings.cis5300-word-embeddings
Word Embeddings and Semantic Similarity (CIS 5300)
Dataset Description
This dataset supports learning about word embeddings — dense vector representations that capture word meaning. It includes a standard similarity benchmark, a word sense disambiguation task, and a Shakespeare corpus for training custom embeddings.
Configs
SimLex-999: Word Similarity Benchmark
SimLex-999 (Hill et al., 2015) is a gold-standard benchmark for evaluating word… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-word-embeddings.malicious-prompts-minilm-embeddingsstargo-embeddings
stargo-embeddings
Dataset repository containing STAR-GO related embedding assets. The metadata.csv is loadable via datasets.load_dataset, while large binaries (e.g. .h5, .npy) are stored as downloadable files.
How to use
Load the metadata table:
from datasets import load_dataset
ds = load_dataset("<your-org-or-username>/<your-dataset-repo>")
print(ds)
Download the large binary assets referenced in the table with hf_hub_download.
HSC-GalaxiesML-VAE-embeddings20newsgroups_embeddings
Dataset Card for feature vector embeddings of the 20newsgroup dataset
Dataset Summary
This dataset contains vector embeddings of the 20newsgroups dataset.
The embeddings were created with the Sentence Transformers library using the multi-qa-MiniLM-L6-cos-v1 model.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data… See the full description on the dataset page: https://huggingface.co/datasets/fscheffczyk/20newsgroups_embeddings.Hugging-Ligand-embeddings
HuggingLigand Dataset
Overview
HuggingLigand is a deep learning pipeline developed to predict the binding affinity between proteins and ligands. This prediction task is essential in fields such as drug discovery, biophysics, and computational biology, where determining how strongly a small molecule ligand binds to a protein target is a key step in understanding molecular interactions and prioritizing drug candidates.
The dataset provides precomputed embeddings for… See the full description on the dataset page: https://huggingface.co/datasets/RSE-Group11/Hugging-Ligand-embeddings.metagenomic_mixture_embeddingswikivoyage-eu-city-embeddings
Dataset Card for Dataset Name
This dataset comprises abstracts from Wikivoyage for 160 European cities along with their corresponding country names, coordinates, and populations. The embeddings are derived from the GTE-Large model, incorporating data from the city, country, population, and abstract columns.
Dataset Sources
Wikivoyage data
World cities database
DepMap_embeddingstaxonomy-embeddings📊 NCBI Dataset
This dataset is derived from the NCBI and was incorporated durgin the pretraining process with CDS-BART.
It contains randonmly selected 500 mRNA sequences from each of the four taxonomies: bacteria, invertebrate, plant, and fungi, totaling 2000 sequences.
⁉️ Dataset Contents
Sequence: The mRNA sequences corresponding to each of the four taxonomies
Label: The labels representing the four different taxonomies: bacteria, invertebrate, plant, and fungi [0,1,2,3]
🎯 Purpose
This… See the full description on the dataset page: https://huggingface.co/datasets/mogam-ai/taxonomy-embeddings.2D_20newsgroups_embeddings
Dataset Card for feature vector embeddings of the 20newsgroup dataset
Dataset Summary
This dataset contains dimensional reduced vector embeddings of the 20newsgroups dataset. This dataset contains two dimensions.
The dimensional reduced embeddings were created with the TruncatedSVD function from the scikit-learn library.
These reduced feature vectors are based on the fscheffczyk/20newsgroup_embeddings dataset.
Supported Tasks and Leaderboards
[More… See the full description on the dataset page: https://huggingface.co/datasets/fscheffczyk/2D_20newsgroups_embeddings.social_security_embeddingsmalicious-prompts-openai-embeddingsArXiv-ML-Title-EmbeddingsThis dataset contains embeddings of the titles of ArXiv Machine Learning papers.
The embeddings are produced from sentence-transformers/paraphrase-MiniLM-L6-v2. The model can be accessed here: HuggingFace Sentence Transformers
The original dataset before embedding can be accessed here: ML ArXiv Papers
faq_embeddingstax-reform-bill-2024-text-embedding-004DinoV2-YGO-card-embeddingsEmbedding_datasetArXiv-ML-Abstract-EmbeddingsThis dataset contains embeddings of the abstracts of ArXiv Machine Learning papers.
The embeddings are produced from sentence-transformers/paraphrase-MiniLM-L6-v2. The model can be accessed here: HuggingFace Sentence Transformers
The original dataset before embedding can be accessed here: ML ArXiv Papers
FAQ_embeddings_examplesp500-business-description-sentence-bert-embeddingsEmbeddings derived from business descriptions of S&P500 companies using sentence-BERT, SentenceTransformer('all-MiniLM-L6-v2') to be exact. For more info on evaluation of sentence transformers (specifcailly the huge GPT-3 versus smaller models see: https://twitter.com/Nils_Reimers/status/1487014195568775173)
sasb_embeddingskaggle-stroke-patients-with-description-embeddingscoraltext-hard-corals-text-traits-for-embedding
CoralText Hard Corals: Text Traits for Embedding
One row per accepted scleractinian (hard / stony) coral species 1,704 species, global scope each carrying a text field built for sentence/document embedding, a set of structured ecological traits and stable identifiers that link back to the source databases. Every row is traceable and the dataset is explicit about where its text comes from and how complete that text is.
This card documents not just what the dataset contains but… See the full description on the dataset page: https://huggingface.co/datasets/xquantize/coraltext-hard-corals-text-traits-for-embedding.jazz-harmony-embeddings
Jazz Harmony Embeddings — 6,900 tune vectors
One 128-dimensional vector per jazz standard, from a small transformer
trained from scratch so that tunes with related harmony — transpositions,
alternate charts, contrafacts — land close together. Produced by the
3-seed ensemble released at
eigenben/jazz-harmony-embeddings;
code and full experiment records at
github.com/eigenben/jazz-harmony-embeddings.
Files
embeddings.npz — embeddings: (6900, 128) float32… See the full description on the dataset page: https://huggingface.co/datasets/eigenben/jazz-harmony-embeddings.yandex-geo-reviews-embeddingsDataset full description: https://www.kaggle.com/datasets/lockiultra/yandex-geo-reviews-embeddings
Dataset contains index column, 768 embedding columns and rating column. Each row corresponds to an embedding representation of the review text with same index.
hf-blogs-jinaai-embeddingsskill_embeddingsembedding-eval-results
Embedding Eval Results
The committed results from a real embedding-model benchmark that embarrassed a leaderboard's recommendation.
What's in here
Four CSV files representing four evaluation runs on a personal Obsidian vault:
File
Notes
Queries
Models
Purpose
20260618-233912.csv
54
29
7
Round 1 — toy slice. Saturated benchmark.
20260619-025330.csv
995
150
7
Stage 1 — real pile. The ranking inverted.
20260620-142445-wholenote.csv
996
450
7
Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/kylebrodeur/embedding-eval-results.
