paragraphs
wikipedia-paragraphs
wikipedia-paragraphs
wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research.
Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID.
Dataset structure
Configurations
The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.openstax_paragraphsTexbooks from openstax.org with their chapters, abstracts and sections.
Sample:
{
"book_title":"World History Volume 1, to 1500",
"language":"en",
"chapters":[
{
"title":"Preface",
"abstract":"None",
"sections":[
{
"title":"About OpenStax",
"paragraph":"OpenStax is part of Rice University, which is a 501(c)(3) nonprofit..."
},
{
"title":"About OpenStax Resources"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/openstax_paragraphs.ccel-paragraphs
CCEL Paragraphs
Dataset Description
Dataset Summary
This dataset includes all paragraphs from the Christian Classics Ethereal Library. It also includes scripture references extracted from the ThML.
Supported Tasks and Leaderboards
It is expected that this dataset can be used as part of the training pipeline for large language models. In particular, it could be used to create a clustering benchmark by using scripture references as labels.… See the full description on the dataset page: https://huggingface.co/datasets/jncraton/ccel-paragraphs.paragraphs-co84b
paragraphs > release-640
https://universe.roboflow.com/roboflow-100/paragraphs-co84b
This dataset is part of RF100, an Intel-sponsored initiative to create a new object detection benchmark for model generalizability.
Dataset Summary
Total images: 6063
Train: 4209 images
Validation: 1221 images
Test: 633 images
Classes: 7 (g, h, g1, g3, -, m, n)
Format: YOLOv8 (Ultralytics)
License: CC BY 4.0
Preprocessing
Auto-orientation of pixel data (with EXIF-orientation… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/paragraphs-co84b.multilingual-wikipedia-paragraphstae-data-split-paragraphs
Split Paragraphs Dataset
Split paragraphs data with configs 000-099.
