umarigan/turkish_clip_dataset_with_text_embeddings
This dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb The main purpose was to extract Turkish captions and download images. You can use this dataset to fine-tune or create a clip model. Since there English and Turkish captions you can also use those to create language model?
1363
This repository belongs to umarigan on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
turkish_clip_dataset_with_text_embeddings
public
creativeml-openrail-m
no
umarigan
