CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1536-1M1M OpenAI Embeddings: text-embedding-3-large 1536 dimensions Created: February 2024. Text used for Embedding: title (string) + text (string) Embedding Model: OpenAI text-embedding-3-large This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here textfeature-extraction1M<n<10M14 likes1.3k downloads3y agoHugging Face02Qdrant /dbpedia-entities-openai3-text-embedding-3-large-3072-1M1M OpenAI Embeddings: text-embedding-3-large 3072 dimensions + ada-002 1536 dimensions — parallel dataset Created: February 2024. Text used for Embedding: title (string) + text (string) Embedding Model: text-embedding-3-large This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here textfeature-extraction1M<n<10M27 likes1k downloads3y agoHugging Face03nielsr /datacomp-small-with-text-embeddings Dataset Card for "datacomp-small-with-text-embeddings" More Information needed image10M<n<100M0 likes957 downloads3y agoHugging Face04umarigan /turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb The main purpose was to extract Turkish captions and download images. You can use this dataset to fine-tune or create a clip model. Since there English and Turkish captions you can also use those to create language model? image100K<n<1M1 likes291 downloads3y agoHugging Face05yoandrey /wiki_text_embeddings Dataset Card for "wiki_text_embeddings" More Information needed text10M<n<100M0 likes218 downloads3y agoHugging Face06filipecosta90 /dbpedia-openai-1M-text-embedding-3-large-3072dtext1M<n<10M1 likes214 downloads1y agoHugging Face07Qdrant /dbpedia-entities-openai3-text-embedding-3-small-1536-100Ktext100K<n<1M7 likes206 downloads3y agoHugging Face08Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1536-100Ktext100K<n<1M2 likes200 downloads3y agoHugging Face09filipecosta90 /dbpedia-openai-1M-text-embedding-3-large-1536dtext1M<n<10M0 likes199 downloads1y agoHugging Face10dhsauojdoiak /h3_stage_text_embedding0 likes167 downloads20d agoHugging Face11AndresR2909 /climate_twitter_text_embeddingstext10K<n<100K0 likes161 downloads2y agoHugging Face12Qdrant /dbpedia-entities-openai3-text-embedding-3-large-3072-100Ktext100K<n<1M2 likes152 downloads3y agoHugging Face13jphwang /twitter_customer_support_weaviate_export_200000_text-embedding-3-smalltext100K<n<1M0 likes150 downloads2y agoHugging Face14distilabel-internal-testing /alvarobartt-improving-text-embeddings-with-llms-full Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml" or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.textn<1K0 likes146 downloads2y agoHugging Face15filipecosta90 /dbpedia-openai-1M-text-embedding-3-large-1024dtext1M<n<10M0 likes133 downloads1y agoHugging Face16sadit /TextEmbeddings0 likes110 downloads2y agoHugging Face17conwaychriscosmo /text-embedding-3-large-english-dictionaryREADME written by Claude inspired by Chris NLTK English Word Embeddings Dataset This dataset contains embeddings for every word in the English language according to the Natural Language Toolkit (NLTK). It provides a comprehensive resource for researchers, developers, and AI enthusiasts working on natural language processing tasks. The dataset is segmented into 7 parts based on alphabetic order. Dataset Overview Source: NLTK English vocabulary Embedding Model: OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/conwaychriscosmo/text-embedding-3-large-english-dictionary.1 likes99 downloads2y agoHugging Face18hllj /synthetic-text-embeddingtexttext-retrieval10K<n<100K0 likes97 downloads2y agoHugging Face19artificial-memory-lab /text-collections-embeddings0 likes90 downloads2y agoHugging Face20filipecosta90 /dbpedia-openai-1M-text-embedding-3-large-512dtext1M<n<10M0 likes81 downloads1y agoHugging Face21Qdrant /dbpedia-entities-openai3-text-embedding-3-small-512-100Ktext100K<n<1M4 likes70 downloads3y agoHugging Face22Qdrant /dbpedia-entities-openai3-text-embedding-3-small-1024-100Ktext100K<n<1M1 likes58 downloads3y agoHugging Face23shayanjm /msmarco__trunc-512__text-embedding-3-largetext100K<n<1M1 likes58 downloads2y agoHugging Face24alvarobartt /improving-text-embeddings-with-llms 🦒 Improving Text Embeddings with Large Language Models Replication of Improving Text Embeddings with Large Language Models. textn<1K7 likes56 downloads3y agoHugging Face25Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1024-100Ktext100K<n<1M2 likes54 downloads3y agoHugging Face26nickmuchi /CFA_Level_1_Text_EmbeddingsVector store of embeddings for CFA Level 1 Curriculum This is a faiss vector store created with Sentence Transformer embeddings using LangChain . Use it for similarity search, question answering or anything else that leverages embeddings! 😃 Creating these embeddings can take a while so here's a convenient, downloadable one 🤗 How to use Download data Load to use with LangChain pip install -qqq langchain sentence_transformers faiss-cpu huggingface_hub import os from langchain.embeddings import… See the full description on the dataset page: https://huggingface.co/datasets/nickmuchi/CFA_Level_1_Text_Embeddings.question-answering3 likes47 downloads3y agoHugging Face27mnlp-nsoai /rag-embeddings-and-texttext100K<n<1M1 likes46 downloads2y agoHugging Face28filipecosta90 /dbpedia-openai-1M-text-embedding-3-large-2048dtext1M<n<10M1 likes45 downloads1y agoHugging Face29DLBDAlkemy /enhanced_reranking_hyde_text-embedding-3-small_queries_with_top5_chunkstext10K<n<100K0 likes38 downloads11mo agoHugging Face30syntropicsignal-ai /wildchat-asking-en-text-embedding-3-small WildChat Asking-mode (EN) — text-embedding-3-small English-language first-turn user prompts from allenai/WildChat-1M, filtered to Asking-mode prompts and embedded with OpenAI's text-embedding-3-small. What's in here Rows 189,916 Language English (en) Embedding model text-embedding-3-small (OpenAI) Embedding dim 1536 (L2-normalized) Format single Parquet file, zstd compression License ODC-BY (inherited from WildChat-1M) Schema… See the full description on the dataset page: https://huggingface.co/datasets/syntropicsignal-ai/wildchat-asking-en-text-embedding-3-small.texttext-retrieval100K<n<1M0 likes36 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.