datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openstax_paragraphsTexbooks from openstax.org with their chapters, abstracts and sections.
Sample:
{
"book_title":"World History Volume 1, to 1500",
"language":"en",
"chapters":[
{
"title":"Preface",
"abstract":"None",
"sections":[
{
"title":"About OpenStax",
"paragraph":"OpenStax is part of Rice University, which is a 501(c)(3) nonprofit..."
},
{
"title":"About OpenStax Resources"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/openstax_paragraphs.wikipedia-paragraphs
Wikipedia Paragraph Samples
Dataset Description
This dataset contains paragraphs extracted from randomly selected English Wikipedia articles. It provides a diverse sample of Wikipedia content across various topics.
Dataset Details
Name: Wikipedia Paragraph Samples
Version: 1.0
Date Created: 2024-08-20
Language: English
Format: JSONLines
Contents
Each line in the dataset represents a single paragraph and contains two fields:
Title of the Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs.wiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.wiki_paragraphs_english
WIKI Paragraphs English
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.wikipedia-paragraphs-complete
Wikipedia Paragraphs Complete Dataset
This dataset consists of English Wikipedia paragraphs ranging from 1 000 to 8 000 characters in length. It was sourced from the Wikimedia dump: "wikimedia/wikipedia", "20231101.en".
Preprocessing Steps
The dataset has undergone extensive cleaning and normalization, including:
Removing brackets
Removing HTML tags
Normalizing bullet points, hyphenated words, quotation marks, Unicode characters, and whitespace
Replacing email… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-complete.grade_labeled_wiki_paragraphs
Grade-Labeled Wiki Paragraphs (GPT-4.1 Nano)
This dataset contains Wikipedia paragraphs simplified to different grade reading levels (targeting Grade 1-12) using the GPT-4.1 Nano model.
Dataset Description
Dataset Summary
The dataset consists of pairs of original Wikipedia paragraphs and their machine-generated simplified versions. The simplification aims to make the text understandable for readers at specific US grade levels while preserving the core… See the full description on the dataset page: https://huggingface.co/datasets/yimingwang123/grade_labeled_wiki_paragraphs.wikipedia-paragraphs-direct-paraphrases
Wikipedia Paragraphs Direct Paraphrases
Paraphrases of Wikipedia paragraphs using AI large language models.
Paragraphs from agentlans/wikipedia-paragraphs-complete sample_k10000 and sample_k50000 splits
Paraphrased using Qwen/Qwen3.5-9B and a distilled Qwen/Qwen3-4B-Instruct-2507 with the following prompt:
Rewrite the following paragraph entirely in your own words while preserving every fact, detail, meaning, nuance, and level of specificity. Do not add, remove, reinterpret… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-direct-paraphrases.lotr_paragraphsopenstax_paragraphsTexbooks from openstax.org with their chapters, abstracts and sections.
Sample:
{
"book_title":"World History Volume 1, to 1500",
"language":"en",
"chapters":[
{
"title":"Preface",
"abstract":"None",
"sections":[
{
"title":"About OpenStax",
"paragraph":"OpenStax is part of Rice University, which is a 501(c)(3) nonprofit..."
},
{
"title":"About OpenStax Resources"… See the full description on the dataset page: https://huggingface.co/datasets/vishwap1991/openstax_paragraphs.michael_paragraphs_units_5
michael_paragraphs_units_5
Dataset uploaded with Python via huggingface_hub.
Files
JSONL source file uploaded to this repository
Notes
Custom dataset
Uploaded automatically from a local file
Synthetic_Clinical_Notes_Paragraphshuivam_finnegans_wake_paragraphs
Dataset Overview
Dataset Name: huivam_finnegans_wake_paragraphs
Creator:
Platform: Hugging Face Datasets
Dataset Context
Source Material: Likely derived from "Finnegans Wake", a novel by James Joyce known for its experimental language and complex structure.
Content Type: Paragraphs (text format)
Expected Dataset Details (Not Selected, Assumed from Typical Dataset Pages)
Key Features:
Text Data: Paragraphs from Finnegans Wake
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/genaforvena/huivam_finnegans_wake_paragraphs.
