datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Estonian-Text-Simplification
Estonian Text Simplification
This repository contains resources and models for Estonian text simplification, including datasets and pre-trained models.
Files
Dataset Files
simplification_training_set.json: A dataset used to fine-tune LLaMA 3.1 for text simplification.
src: The source of the data.
original: The original sentence to be simplified.
simpl_lex: A lexical simplification (may be empty).
simpl_final: The final simplified sentence.… See the full description on the dataset page: https://huggingface.co/datasets/vulturuldemare/Estonian-Text-Simplification.age-specific-text-simplification
Age-Specific Text Simplification Dataset
Dataset Description
This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group.
Dataset Summary
Total Examples: 17,177
Training Split: 15,459 examples
Validation Split: 1,718 examples
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.nemotron-cc-atomic-simplification-gemma4-31b
nemotron-cc atomic-statement simplification (Gemma 4 31B-it)
2,000,000 records: source text from
nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic
split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object
statements.
Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the
producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to
<=8192 templated tokens.
Fields
id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.Leesplank_NL_wikipedia_simplificationsThe set contains 2.87M pragraphs of prompt/result combinations, where the prompt is a paragraph from Dutch Wikipedia and the result is a simplified text, which could include more than one paragraph.
This dataset was created by UWV, as a part of project "Leesplank", an effort to generate datasets that are ethically and legally sound.
The basis of this dataset was the wikipedia extract as a part of Gigacorpus (http://gigacorpus.nl/). The lines were fed one by one into GPT 4 1106 preview, where… See the full description on the dataset page: https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications.spanish_simplification_20k
Dataset Card: spanish_simplification_20k
This dataset provides 20K samples for Spanish simplification across a diverse range of topics, designed for people with cognitive challenges. The samples vary in length, from short to long. This dataset can supplement a larger training corpus for fine-tuning small language models, such as Gemma 4B, to simplify complex Spanish into simpler Spanish or translate English directly into simplified Spanish.
Dataset Contributors… See the full description on the dataset page: https://huggingface.co/datasets/khaledmahmoud/spanish_simplification_20k.medical-reports-simplification-dataset
🏥 Medical Reports Simplification Dataset
📋 Description
Dataset créé avec Gemini 2.5 Pro (Preview) pour entraîner des modèles à simplifier les rapports médicaux complexes en explications compréhensibles pour les patients.
🎯 Objectif : Démocratiser l'accès à l'information médicale en rendant les rapports techniques accessibles au grand public.
🔧 Génération du Dataset
Génération : Gemini 2.5 Pro (Preview)
Validation : Contrôle qualité automatisé… See the full description on the dataset page: https://huggingface.co/datasets/Sadou/medical-reports-simplification-dataset.adaption-legal-clause-simplification
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-legal_clause_simplification
This dataset contains pairs of formal legal or policy clauses and their simplified, plain-language equivalents. The prompts feature complex contractual language regarding payments, confidentiality, and benefits, while the completions provide clear, concise restatements. It is designed for training models to translate legalese into accessible text.… See the full description on the dataset page: https://huggingface.co/datasets/Karthikrv/adaption-legal-clause-simplification.adaption-legal-text-simplification
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-legal_text_simplification
This dataset contains pairs of formal legal or institutional text prompts and their corresponding simplified, plain-language completions. The content focuses on translating complex terms regarding consumer responsibilities, examination policies, and liability into clear, everyday English. It is designed for training models to perform legal text… See the full description on the dataset page: https://huggingface.co/datasets/Karthikrv/adaption-legal-text-simplification.
