dyslexia
Datasets
All datasets matching “dyslexia”wmt14_injected_synthetic_dyslexia
Dataset Summary
The WMT14 injected synthetic dyslexia dataset is a modified version of the WMT14 English test set. This dataset was created to test the capabilities of SOTA machine translations models on dyslexic style text. This research was supported by AImpower.org.
How the data is structured
In "Data/French_translated_data", each file within the dataset consists of a “.txt” or “.docx” file containing the translated sentences from AWS, Google, Azure and OpenAI.
In… See the full description on the dataset page: https://huggingface.co/datasets/gpric024/wmt14_injected_synthetic_dyslexia.dyslexia-french-cp-ce1
FrenchDyslexia-CP-CE1 — French Dyslexia Text Simplification Dataset
Dataset Description
A manually annotated parallel corpus of 362 French text pairs for dyslexia-adapted text simplification, targeting children aged 6–8 (CP and CE1) in Moroccan French-medium primary schools.
Each pair consists of an original French primary school text and its phonics-aware simplified version, produced according to a 13-rule simplification referential including:
Short sentences… See the full description on the dataset page: https://huggingface.co/datasets/lsadouk1111/dyslexia-french-cp-ce1.akura-sinhala-dyslexia-datasetakura-dyslexia-sinhala
🧠 Akura AI - Sinhala Dyslexia Correction Dataset
This dataset contains examples of Sinhala text with common dyslexic writing errors and their corrections. It is designed for fine-tuning LLMs (like Llama 3) to detect and correct these specific patterns.
Dataset Structure
The dataset is in JSONL format, suitable for chat-based model fine-tuning.
Example Entry
{
"messages": [
{"role": "system", "content": "You are Akura AI. Analyze the Sinhala text for… See the full description on the dataset page: https://huggingface.co/datasets/hasinduOnline/akura-dyslexia-sinhala.dyslexia-handwriting-datasetdyslexia-simplified-text-dataset
