datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bfs-optimal-symbolic-algebra-simplification-trajectories
ISRE v7 BFS Trajectories
This dataset contains BFS-optimal symbolic algebra simplification trajectories
for the ISRE research project.
Each row is one trajectory. Nested fields are stored as JSON strings to preserve
the original AST and step structure without lossy flattening.
What This Dataset Is
This is not a natural-language instruction dataset. It is a symbolic algebra
policy-learning dataset.
Each trajectory starts from a deliberately scrambled algebraic… See the full description on the dataset page: https://huggingface.co/datasets/NecroMOnk/bfs-optimal-symbolic-algebra-simplification-trajectories.samer-arabic-text-simplification
SAMER Arabic Text Simplification Dataset (Cleaned Version)
Description
This dataset is a cleaned and structured version of the SAMER Corpus (The SAMER Arabic Text Simplification Corpus). It is prepared specifically for training Seq2Seq models (e.g., AraT5) and fine-tuning Large Language Models (LLMs) on Arabic Text Simplification and Readability Assessment tasks.
Dataset Structure
The dataset contains the following fields:
clean_text: Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/vn3er/samer-arabic-text-simplification.samer-arabic-simplification-onlylegal-simplification-pt
Dataset Structure
Configuration: random_split
train_random
validation_random
test_random
Configuration: by_source
acordaos_tcu
stf_decisions
stf_votes
tjsp
trf5
ulysses_tesemo
