CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nielsr /datacomp-small-with-text-embeddings Dataset Card for "datacomp-small-with-text-embeddings" More Information needed image10M<n<100M0 likes957 downloads3y agoHugging Face02umarigan /turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb The main purpose was to extract Turkish captions and download images. You can use this dataset to fine-tune or create a clip model. Since there English and Turkish captions you can also use those to create language model? image100K<n<1M1 likes291 downloads3y agoHugging Face03yoandrey /wiki_text_embeddings Dataset Card for "wiki_text_embeddings" More Information needed text10M<n<100M0 likes218 downloads3y agoHugging Face04AndresR2909 /climate_twitter_text_embeddingstext10K<n<100K0 likes161 downloads2y agoHugging Face05distilabel-internal-testing /alvarobartt-improving-text-embeddings-with-llms-full Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml" or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.textn<1K0 likes146 downloads2y agoHugging Face06sadit /TextEmbeddings0 likes110 downloads2y agoHugging Face07artificial-memory-lab /text-collections-embeddings0 likes90 downloads2y agoHugging Face08alvarobartt /improving-text-embeddings-with-llms 🦒 Improving Text Embeddings with Large Language Models Replication of Improving Text Embeddings with Large Language Models. textn<1K7 likes56 downloads3y agoHugging Face09nickmuchi /CFA_Level_1_Text_EmbeddingsVector store of embeddings for CFA Level 1 Curriculum This is a faiss vector store created with Sentence Transformer embeddings using LangChain . Use it for similarity search, question answering or anything else that leverages embeddings! 😃 Creating these embeddings can take a while so here's a convenient, downloadable one 🤗 How to use Download data Load to use with LangChain pip install -qqq langchain sentence_transformers faiss-cpu huggingface_hub import os from langchain.embeddings import… See the full description on the dataset page: https://huggingface.co/datasets/nickmuchi/CFA_Level_1_Text_Embeddings.question-answering3 likes47 downloads3y agoHugging Face10mnlp-nsoai /rag-embeddings-and-texttext100K<n<1M1 likes46 downloads2y agoHugging Face11uzair921 /SKILLSPAN_embeddings_texttext1K<n<10K0 likes33 downloads2y agoHugging Face12LGirrbach /fg-clip-text-embeddingstext10K<n<100K0 likes26 downloads20d agoHugging Face13biomedical-translator /monarch_kg_node_text_embeddingstabular1M<n<10M0 likes25 downloads2y agoHugging Face14Sreenath /million-text-embeddings Million Text Embeddings A dataset with more than a million English sentences and their respective embeddings with the all-mpnet-base-v2 model.Train Set: 1,000,000Test Set: 2,00,000Dimensions: 768Source: agentlans/high-quality-english-sentences GitHub: sreenaths/hf-datasets text1M<n<10M1 likes21 downloads2y agoHugging Face15biomedical-translator /maxo-text-embeddingstext10K<n<100K0 likes19 downloads2y agoHugging Face16kvriza8 /clip_microscopy_image_text_embeddingstext10K<n<100K0 likes19 downloads2y agoHugging Face17uzair921 /wnut_17_embeddings_texttext1K<n<10K0 likes18 downloads2y agoHugging Face18uzair921 /conll2003_embeddings_texttext10K<n<100K0 likes17 downloads2y agoHugging Face19seungheondoh /eval-music-text-embeddingstext1K<n<10K0 likes17 downloads1y agoHugging Face20biomedical-translator /bfo-text-embeddingstextn<1K0 likes16 downloads2y agoHugging Face21biomedical-translator /go-text-embeddingstext10K<n<100K0 likes13 downloads2y agoHugging Face22distilabel-internal-testing /alvarobartt-improving-text-embeddings-with-llms Dataset Card for alvarobartt-improving-text-embeddings-with-llms This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms/raw/main/pipeline.yaml" or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms.textn<1K0 likes10 downloads2y agoHugging Face23rk404 /text_embeddings0 likes9 downloads3y agoHugging Face24biomedical-translator /chebi-text-embeddingstextn<1K0 likes9 downloads2y agoHugging Face25jpearce610 /text_embeddingstabular1M<n<10M0 likes9 downloads8mo agoHugging Face26biomedical-translator /mondo-text-embeddingstext10K<n<100K0 likes8 downloads2y agoHugging Face27davidberenstein1957 /hf-blogs-text-embeddingstext1K<n<10K0 likes8 downloads2y agoHugging Face28MikeGreen2710 /1m7_remote_text_embeddingstext1M<n<10M0 likes7 downloads1y agoHugging Face29uzair921 /QWEN_CONLL2003_EMBEDDINGS_TEXT_LLM_RAG_25_openai_Texttext10K<n<100K0 likes6 downloads2y agoHugging Face30uzair921 /QWEN_CONLL2003_EMBEDDINGS_TEXT_LLM_RAG_75_openai_Texttext10K<n<100K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.