datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb
The main purpose was to extract Turkish captions and download images.
You can use this dataset to fine-tune or create a clip model.
Since there English and Turkish captions you can also use those to create language model?
wiki_text_embeddings
Dataset Card for "wiki_text_embeddings"
More Information needed
climate_twitter_text_embeddingsalvarobartt-improving-text-embeddings-with-llms-full
Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.TextEmbeddingstext-collections-embeddingsimproving-text-embeddings-with-llms
🦒 Improving Text Embeddings with Large Language Models
Replication of Improving Text Embeddings with Large Language Models.
CFA_Level_1_Text_EmbeddingsVector store of embeddings for CFA Level 1 Curriculum
This is a faiss vector store created with Sentence Transformer embeddings using LangChain . Use it for similarity search, question answering or anything else that leverages embeddings! 😃
Creating these embeddings can take a while so here's a convenient, downloadable one 🤗
How to use
Download data
Load to use with LangChain
pip install -qqq langchain sentence_transformers faiss-cpu huggingface_hub
import os
from langchain.embeddings import… See the full description on the dataset page: https://huggingface.co/datasets/nickmuchi/CFA_Level_1_Text_Embeddings.rag-embeddings-and-textSKILLSPAN_embeddings_textfg-clip-text-embeddingsmonarch_kg_node_text_embeddingsmillion-text-embeddings
Million Text Embeddings
A dataset with more than a million English sentences and their respective embeddings with the all-mpnet-base-v2 model.Train Set: 1,000,000Test Set: 2,00,000Dimensions: 768Source: agentlans/high-quality-english-sentences GitHub: sreenaths/hf-datasets
maxo-text-embeddingsclip_microscopy_image_text_embeddingswnut_17_embeddings_textconll2003_embeddings_texteval-music-text-embeddingsbfo-text-embeddingsgo-text-embeddingsalvarobartt-improving-text-embeddings-with-llms
Dataset Card for alvarobartt-improving-text-embeddings-with-llms
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms.text_embeddingschebi-text-embeddingstext_embeddingsmondo-text-embeddingshf-blogs-text-embeddings1m7_remote_text_embeddingsQWEN_CONLL2003_EMBEDDINGS_TEXT_LLM_RAG_25_openai_TextQWEN_CONLL2003_EMBEDDINGS_TEXT_LLM_RAG_75_openai_Text
