datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
samer-arabic-text-simplification
SAMER Arabic Text Simplification Dataset (Cleaned Version)
Description
This dataset is a cleaned and structured version of the SAMER Corpus (The SAMER Arabic Text Simplification Corpus). It is prepared specifically for training Seq2Seq models (e.g., AraT5) and fine-tuning Large Language Models (LLMs) on Arabic Text Simplification and Readability Assessment tasks.
Dataset Structure
The dataset contains the following fields:
clean_text: Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/vn3er/samer-arabic-text-simplification.TextSimplification-PTBR-330kenglish-text-simplification-for-finetuninghindi-text-simplification-for-finetuningawadhi-text-simplification-for-finetuning
