CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opendatalab /SlimPajama-Meta-rater-Readability-30B Top 30B token SlimPajama Subset selected by the Readability rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.tabulartext-generation1M<n<10M1 likes358 downloads1y agoHugging Face02somosnlp-hackathon-2022 /readability-es-hackathon-pln-public Dataset Card for [readability-es-sentences] Dataset Description Compilation of short Spanish articles for readability assessment. Dataset Summary This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources: Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.texttext-classification1K<n<10K3 likes124 downloads3y agoHugging Face03agentlans /low-readability-text Low Readability Text Dataset This dataset consists of high-complexity English web text with an estimated readability at or above the U.S. Grade 12 level. The content typically features advanced, highly technical prose or verbose syntactical structures, making it well-suited for researching complex language understanding and automation. Primary Use Cases Text Simplification: Training and evaluating models to translate complex text into plain English. Information… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/low-readability-text.texttext-generation100K<n<1M0 likes49 downloads4mo agoHugging Face04somosnlp-hackathon-2022 /readability-es-caes Dataset Card for [readability-es-caes] Dataset Description Dataset Summary This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources: CAES corpus (Martínez et al., 2019): the "Corpus de Aprendices del Español" is a collection of texts produced by Spanish L2 learners from Spanish learning centers and universities. These text are produced by students… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-caes.texttext-classification10K<n<100K3 likes48 downloads3y agoHugging Face05agentlans /advanced-readability-analysis Advanced Readability Analysis This dataset provides rich syntactic and lexical complexity features calculated from English text snippets. It is designed to help researchers study the underlying factors that influence reading difficulty, especially in cases where traditional readability formulas yield conflicting results. The source text is pulled from the training split of the agentlans/readability dataset. The linguistic annotations and complexity metrics were computed using a… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/advanced-readability-analysis.tabularfeature-extraction10K<n<100K1 likes6 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.