datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dbpedia-entities-openai3-text-embedding-3-large-1536-1M1M OpenAI Embeddings: text-embedding-3-large 1536 dimensions
Created: February 2024.
Text used for Embedding: title (string) + text (string)
Embedding Model: OpenAI text-embedding-3-large
This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here
dbpedia-entities-openai3-text-embedding-3-large-3072-1M1M OpenAI Embeddings: text-embedding-3-large 3072 dimensions + ada-002 1536 dimensions — parallel dataset
Created: February 2024.
Text used for Embedding: title (string) + text (string)
Embedding Model: text-embedding-3-large
This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here
datacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb
The main purpose was to extract Turkish captions and download images.
You can use this dataset to fine-tune or create a clip model.
Since there English and Turkish captions you can also use those to create language model?
wiki_text_embeddings
Dataset Card for "wiki_text_embeddings"
More Information needed
dbpedia-openai-1M-text-embedding-3-large-3072ddbpedia-entities-openai3-text-embedding-3-small-1536-100Kdbpedia-entities-openai3-text-embedding-3-large-1536-100Kdbpedia-openai-1M-text-embedding-3-large-1536dh3_stage_text_embeddingclimate_twitter_text_embeddingsdbpedia-entities-openai3-text-embedding-3-large-3072-100Ktwitter_customer_support_weaviate_export_200000_text-embedding-3-smallalvarobartt-improving-text-embeddings-with-llms-full
Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.dbpedia-openai-1M-text-embedding-3-large-1024dTextEmbeddingstext-embedding-3-large-english-dictionaryREADME written by Claude inspired by Chris
NLTK English Word Embeddings Dataset
This dataset contains embeddings for every word in the English language according to the Natural Language Toolkit (NLTK).
It provides a comprehensive resource for researchers, developers, and AI enthusiasts working on natural language processing tasks.
The dataset is segmented into 7 parts based on alphabetic order.
Dataset Overview
Source: NLTK English vocabulary
Embedding Model: OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/conwaychriscosmo/text-embedding-3-large-english-dictionary.synthetic-text-embeddingtext-collections-embeddingsdbpedia-openai-1M-text-embedding-3-large-512ddbpedia-entities-openai3-text-embedding-3-small-512-100Kdbpedia-entities-openai3-text-embedding-3-small-1024-100Kmsmarco__trunc-512__text-embedding-3-largeimproving-text-embeddings-with-llms
🦒 Improving Text Embeddings with Large Language Models
Replication of Improving Text Embeddings with Large Language Models.
dbpedia-entities-openai3-text-embedding-3-large-1024-100KCFA_Level_1_Text_EmbeddingsVector store of embeddings for CFA Level 1 Curriculum
This is a faiss vector store created with Sentence Transformer embeddings using LangChain . Use it for similarity search, question answering or anything else that leverages embeddings! 😃
Creating these embeddings can take a while so here's a convenient, downloadable one 🤗
How to use
Download data
Load to use with LangChain
pip install -qqq langchain sentence_transformers faiss-cpu huggingface_hub
import os
from langchain.embeddings import… See the full description on the dataset page: https://huggingface.co/datasets/nickmuchi/CFA_Level_1_Text_Embeddings.rag-embeddings-and-textdbpedia-openai-1M-text-embedding-3-large-2048denhanced_reranking_hyde_text-embedding-3-small_queries_with_top5_chunkswildchat-asking-en-text-embedding-3-small
WildChat Asking-mode (EN) — text-embedding-3-small
English-language first-turn user prompts from
allenai/WildChat-1M,
filtered to Asking-mode prompts and embedded with OpenAI's
text-embedding-3-small.
What's in here
Rows
189,916
Language
English (en)
Embedding model
text-embedding-3-small (OpenAI)
Embedding dim
1536 (L2-normalized)
Format
single Parquet file, zstd compression
License
ODC-BY (inherited from WildChat-1M)
Schema… See the full description on the dataset page: https://huggingface.co/datasets/syntropicsignal-ai/wildchat-asking-en-text-embedding-3-small.
