text-embeddings
jina-embeddings-v5-text-smalljina-embeddings-v5-text-nano-retrievaljina-embeddings-v5-text-nanojina-embeddings-v5-text-small-retrievaljina-embeddings-v5-text-small-text-matchingjina-embeddings-v5-text-nano-retrieval-GGUFjina-embeddings-v5-text-small-classificationjina-embeddings-v5-text-nano-text-matching
datacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb
The main purpose was to extract Turkish captions and download images.
You can use this dataset to fine-tune or create a clip model.
Since there English and Turkish captions you can also use those to create language model?
wiki_text_embeddings
Dataset Card for "wiki_text_embeddings"
More Information needed
climate_twitter_text_embeddingsalvarobartt-improving-text-embeddings-with-llms-full
Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.TextEmbeddings
