datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationArabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240This dataset repo contains the dataset (CC3M+CC12M+SBU) translated using opus-mt-en-ar and cleaned. Its size about 13M
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({
train: Dataset({
features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'],
num_rows: 12166802
})
})
ccs_synthetic_ar_1M-Arabic_dataset_1M_translated_jsonl_formatarabic_dataset_translated_v2_ViT-B-16-SigLIP-512
