CoolFace
Datasetpublic

Sreenath/million-text-embeddings

Million Text Embeddings A dataset with more than a million English sentences and their respective embeddings with the all-mpnet-base-v2 model.Train Set: 1,000,000Test Set: 2,00,000Dimensions: 768Source: agentlans/high-quality-english-sentences GitHub: sreenaths/hf-datasets

sourceHugging Faceodc-byupdated 2y agoView on Hugging Face
1likes21downloads
Dataset Card

Million Text Embeddings

A dataset with more than a million English sentences and their respective embeddings with the all-mpnet-base-v2 model.\ Train Set: 1,000,000\ Test Set: 2,00,000\ Dimensions: 768\ Source: [agentlans/high-quality-english-sentences](https://huggingface.co/datasets/agentlans/high-quality-english-sentences) \ GitHub: [sreenaths/hf-datasets](https://github.com/sreenaths/hf-datasets/blob/main/million-text-embeddings.ipynb)