datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.wiki_paragraphs_english
WIKI Paragraphs English
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.grade_labeled_wiki_paragraphs
Grade-Labeled Wiki Paragraphs (GPT-4.1 Nano)
This dataset contains Wikipedia paragraphs simplified to different grade reading levels (targeting Grade 1-12) using the GPT-4.1 Nano model.
Dataset Description
Dataset Summary
The dataset consists of pairs of original Wikipedia paragraphs and their machine-generated simplified versions. The simplification aims to make the text understandable for readers at specific US grade levels while preserving the core… See the full description on the dataset page: https://huggingface.co/datasets/yimingwang123/grade_labeled_wiki_paragraphs.michael_paragraphs_units_5
michael_paragraphs_units_5
Dataset uploaded with Python via huggingface_hub.
Files
JSONL source file uploaded to this repository
Notes
Custom dataset
Uploaded automatically from a local file
