datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimPajama-Meta-rater-Readability-30B
Top 30B token SlimPajama Subset selected by the Readability rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.CommonLit-Ease-of-ReadabilityFinRAD_Financial_Readability_Assessment_Dataset
FinRAD: Financial Readability Assessment Dataset - 13,000+ Definitions of Financial Terms for Measuring Readability
This repository contains the dataset mentioned in the paper: FinRAD: Financial Readability Assessment Dataset - 13,000+ Definitions of Financial Terms for Measuring Readability (presented at The Financial Narrative Processing Workshop colocated with LREC-2022, Marseille, France).
In addition to this, data collection & cleaning scripts, embedding extraction & model… See the full description on the dataset page: https://huggingface.co/datasets/sohomghosh/FinRAD_Financial_Readability_Assessment_Dataset.oncology-readability-collapse-risk-v0.3
What this dataset does
This dataset tests whether a model can detect pre-cancer instability risk from loss of signal readability rather than from stress burden alone.
The task is not cancer diagnosis.
The task is to classify whether a synthetic tissue ecology has entered readability collapse risk.
Core Stability Idea
The dataset represents a stability-transition hypothesis.
Cancer vulnerability may begin when tissue regulation loses the ability to correctly read… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/oncology-readability-collapse-risk-v0.3.Dorn_Code_Readability
Software Readability Dataset
This repository contains the dataset used to build and evaluate the readability model presented in:
A General Software Readability Model
Jonathan Dorn & Westley Weimer, University of Virginia
The dataset consists of human-annotated code snippets sampled from real open-source projects and labeled for perceived readability. It is the largest such dataset collected for software readability research to date.
📦 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/rufimelo/Dorn_Code_Readability.advanced-readability-analysis
Advanced Readability Analysis
This dataset provides rich syntactic and lexical complexity features calculated from English text snippets. It is designed to help researchers study the underlying factors that influence reading difficulty, especially in cases where traditional readability formulas yield conflicting results.
The source text is pulled from the training split of the agentlans/readability dataset.
The linguistic annotations and complexity metrics were computed using a… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/advanced-readability-analysis.aozorabunko_readability_score
