datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
low-quality-multilingual-sentences
Low Quality Multilingual Sentences
This dataset is a complement to agentlans/high-quality-multilingual-sentences to extend it to more languages.
The new sentences in this dataset are low quality, proceed with caution.
low-readability-text
Low Readability Text Dataset
This dataset consists of high-complexity English web text with an estimated readability at or above the U.S. Grade 12 level. The content typically features advanced, highly technical prose or verbose syntactical structures, making it well-suited for researching complex language understanding and automation.
Primary Use Cases
Text Simplification: Training and evaluating models to translate complex text into plain English.
Information… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/low-readability-text.
