arabic-si
ArabicNER-Wojoodbert-base-arabertv2-ViT-B-16-SigLIP-512-epoch-155-trained-2MArabicWojood-FlatNERbert-base-arabic-camelbert-msa-sixteenthbert-base-arabic-camelbert-msa-sixteenth-xnli-finetunedbert-base-arabertv2-ViT-B-16-SigLIP-512-epoch-55-trained-2Mbert-base-arabertv2-ViT-B-16-SigLIP-512-epoch-155-trained-2M-fp32-this-correctarat5-arabic-simplification
karsl-502-arabic-sign-language-v2silma-arabic-english-sts-dataset-v1.0
SILMA STS Arabic/English Dataset - v1.0
Overview
The SILMA STS Arabic/English Dataset - v1.0 is a dataset designed for training and evaluating sentence embeddings for Arabic and English tasks. It consists of five different splits that cover monolingual and multilingual sentence pairs, with human-annotated similarity scores. The dataset includes both Arabic-to-Arabic and English-to-English pairs, as well as cross-lingual Arabic-English pairs, making it a valuable resource… See the full description on the dataset page: https://huggingface.co/datasets/silma-ai/silma-arabic-english-sts-dataset-v1.0.karsl-502-arabic-sign-language-finalsitr-arabic-pii
سِتر · Sitr
An Arabic-first PII detection dataset for Saudi Arabia, the Gulf and the Levant
Character-level spans for 43 personal-data types across 17,899 Arabic examples — regional dialects, code-switching, call-centre transcripts, and the long documents people paste into AI assistants.
Why this exists
General-purpose PII filters underperform badly in Arabic. Independent evaluation of the widely used openai/privacy-filter reports… See the full description on the dataset page: https://huggingface.co/datasets/mabahboh/sitr-arabic-pii.samer-arabic-text-simplification
SAMER Arabic Text Simplification Dataset (Cleaned Version)
Description
This dataset is a cleaned and structured version of the SAMER Corpus (The SAMER Arabic Text Simplification Corpus). It is prepared specifically for training Seq2Seq models (e.g., AraT5) and fine-tuning Large Language Models (LLMs) on Arabic Text Simplification and Readability Assessment tasks.
Dataset Structure
The dataset contains the following fields:
clean_text: Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/vn3er/samer-arabic-text-simplification.arabic-triplets-1m-curated-sims-len
Arabic 1 Million Triplets (curated):
This is a curated dataset to use in Arabic ColBERT and SBERT models (among other uses).
In addition to anchor, positive and negative columns, the dataset has two columns: sim_pos and sim_neg which are cosine
similarities between the anchor (query) and bothe positive and negative examples.The last 3 columns are lengths (words) for each of the anchor, positive and negative examples. Length uses simple split on space, not tokens.
The cosine… See the full description on the dataset page: https://huggingface.co/datasets/akhooli/arabic-triplets-1m-curated-sims-len.
