datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dbpedia-entities-openai3-text-embedding-3-large-1536-1M1M OpenAI Embeddings: text-embedding-3-large 1536 dimensions
Created: February 2024.
Text used for Embedding: title (string) + text (string)
Embedding Model: OpenAI text-embedding-3-large
This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here
dbpedia-entities-openai3-text-embedding-3-large-3072-1M1M OpenAI Embeddings: text-embedding-3-large 3072 dimensions + ada-002 1536 dimensions — parallel dataset
Created: February 2024.
Text used for Embedding: title (string) + text (string)
Embedding Model: text-embedding-3-large
This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here
datacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb
The main purpose was to extract Turkish captions and download images.
You can use this dataset to fine-tune or create a clip model.
Since there English and Turkish captions you can also use those to create language model?
wiki_text_embeddings
Dataset Card for "wiki_text_embeddings"
More Information needed
dbpedia-openai-1M-text-embedding-3-large-3072ddbpedia-entities-openai3-text-embedding-3-small-1536-100Kdbpedia-entities-openai3-text-embedding-3-large-1536-100Kdbpedia-openai-1M-text-embedding-3-large-1536dclimate_twitter_text_embeddingsdbpedia-entities-openai3-text-embedding-3-large-3072-100Kalvarobartt-improving-text-embeddings-with-llms-full
Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.dbpedia-openai-1M-text-embedding-3-large-1024dtwitter_customer_support_weaviate_export_200000_text-embedding-3-smallsynthetic-text-embeddingdbpedia-openai-1M-text-embedding-3-large-512ddbpedia-entities-openai3-text-embedding-3-small-512-100Kdbpedia-entities-openai3-text-embedding-3-small-1024-100Kmsmarco__trunc-512__text-embedding-3-largeimproving-text-embeddings-with-llms
🦒 Improving Text Embeddings with Large Language Models
Replication of Improving Text Embeddings with Large Language Models.
dbpedia-entities-openai3-text-embedding-3-large-1024-100Krag-embeddings-and-textdbpedia-openai-1M-text-embedding-3-large-2048denhanced_reranking_hyde_text-embedding-3-small_queries_with_top5_chunkswildchat-asking-en-text-embedding-3-small
WildChat Asking-mode (EN) — text-embedding-3-small
English-language first-turn user prompts from
allenai/WildChat-1M,
filtered to Asking-mode prompts and embedded with OpenAI's
text-embedding-3-small.
What's in here
Rows
189,916
Language
English (en)
Embedding model
text-embedding-3-small (OpenAI)
Embedding dim
1536 (L2-normalized)
Format
single Parquet file, zstd compression
License
ODC-BY (inherited from WildChat-1M)
Schema… See the full description on the dataset page: https://huggingface.co/datasets/syntropicsignal-ai/wildchat-asking-en-text-embedding-3-small.SKILLSPAN_embeddings_textenhanced_text-embedding-3-small_queries_with_top5_chunkstext-embedding-dataset
Text embedding Datasets
The text embedding datasets consist of several (query, passage) paired datasets aiming for text-embedding model finetuning. These datasets are ideal for developing and testing algorithms in the fields of natural language processing, information retrieval, and similar applications.
Dataset Details
Each dataset in this collection is structured to facilitate the training and evaluation of text-embedding models. The datasets are diverse, covering… See the full description on the dataset page: https://huggingface.co/datasets/ProfessorBob/text-embedding-dataset.legal-text-embeddingenhanced_text-embedding-3-small_queries_with_top5_chunks_answers_Qwen3-0.6B
