CoolFace
20 results

readability

penfever /meta-llama_Llama-3.1-8B-Instruct-jdgfct-Readabilitytext100K<n<1M0 likes368 downloads5mo agoHugging Faceopendatalab /SlimPajama-Meta-rater-Readability-30B Top 30B token SlimPajama Subset selected by the Readability rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.tabulartext-generation1M<n<10M1 likes356 downloads1y agoHugging Facepenfever /nvidia_NVLM-D-72B-jdgfct-Readabilitytext100K<n<1M0 likes251 downloads5mo agoHugging Facepenfever /meta-llama_Llama-3.1-70B-Instruct-jdgfct-Readabilitytext100K<n<1M0 likes249 downloads2y agoHugging Facesomosnlp-hackathon-2022 /readability-es-hackathon-pln-public Dataset Card for [readability-es-sentences] Dataset Description Compilation of short Spanish articles for readability assessment. Dataset Summary This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources: Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.texttext-classification1K<n<10K3 likes128 downloads3y agoHugging Faceagentlans /low-readability-text Low Readability Text Dataset This dataset consists of high-complexity English web text with an estimated readability at or above the U.S. Grade 12 level. The content typically features advanced, highly technical prose or verbose syntactical structures, making it well-suited for researching complex language understanding and automation. Primary Use Cases Text Simplification: Training and evaluating models to translate complex text into plain English. Information… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/low-readability-text.texttext-generation100K<n<1M0 likes99 downloads4mo agoHugging Face