datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-paragraphs
wikipedia-paragraphs
wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research.
Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID.
Dataset structure
Configurations
The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the
exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a single paragraph of dense
reasoning about the immediate continuation. Only the part after </think> — the paragraph —
becomes the VALUE. The reasoning inside the think block is dropped.
The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.d1_code_long_paragraphswiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.japanese-wikipedia-paragraphsA slightly modified version of the parsing and chunking method for singletongue/wikipedia-utils.
Pre-processing was performed using oshizo/wikipedia-utils, which is a fork of the original repository, singletongue/wikipedia-utils.
The Wikipedia data was crawled between 2023/12/5 and 2023/12/8.
books3_lowgrade_paragraphs
Dataset Card for "books3_lowgrade_paragraphs"
the_pile books3, books with smog grade difficulty estimate between 6.6 or and 7.1. Split into paragraphs and filtered out most 'non-paragraphs' like titles, tables of content, etc.
For easier books, see books3_basic_paragraphs
arxiv-paragraphsGutenQA_Paragraphs
📚 GutenQA-Paragraphs
GutenQA-Paragraphs consists on the same 100 Public Domain Narrative Books used in GutenQA. In this version, passages are extracted at the paragraph level.
The GutenQA dataset, as available, is the result of applying the text segmentation method LumberChunker to the GutenQA-Paragraphs.
The dataset is organized into the following columns:
Book Name: The title of the book from which the passage is extracted.
Book ID: A unique integer identifier assigned to each… See the full description on the dataset page: https://huggingface.co/datasets/LumberChunker/GutenQA_Paragraphs.wiki_paragraphs_english
WIKI Paragraphs English
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3
Dataset Card for "jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3"
More Information needed
d1_math_long_paragraphsResearchArcade-arxiv-paragraphsbooks3_basic_paragraphs
Dataset Card for "books3_basic_paragraphs"
the_pile books3, books with smog grade difficulty estimate of 6.5 or under. Split into paragraphs and filtered out most 'non-paragraphs' like titles, tables of content, etc.
paragraphs
Dataset Card for "paragraphs"
More Information needed
gp-long-paragraphsgrade_labeled_wiki_paragraphs
Grade-Labeled Wiki Paragraphs (GPT-4.1 Nano)
This dataset contains Wikipedia paragraphs simplified to different grade reading levels (targeting Grade 1-12) using the GPT-4.1 Nano model.
Dataset Description
Dataset Summary
The dataset consists of pairs of original Wikipedia paragraphs and their machine-generated simplified versions. The simplification aims to make the text understandable for readers at specific US grade levels while preserving the core… See the full description on the dataset page: https://huggingface.co/datasets/yimingwang123/grade_labeled_wiki_paragraphs.arxiv-paragraph-tablesRuBQ_2.0-paragraphs
RuBQ_2.0-paragraphs
For test and dev data see d0rj/RuBQ_2.0
jawiki-paragraphs
Dataset Card for llm-book/jawiki-paragraphs
書籍『大規模言語モデル入門』で使用する Wikipedia 段落のデータセットです。
GitHub リポジトリ singletongue/wikipedia-utils で公開されているデータセットを利用しています。
Licence
本データセットで使用している Wikipedia のコンテンツは、クリエイティブ・コモンズ表示・継承ライセンス 3.0 (CC BY-SA 3.0) および GNU 自由文書ライセンス (GFDL) の下に配布されているものです。
d1_science_long_paragraphsd1_code_long_paragraphs_10kResearchArcade-arxiv-paragraph-figuresarxiv-paragraph-figuresResearchArcade-arxiv-paragraph-citationssecrethackatondata_paragraphsBundestagsreden_Paragraphs_with_likes
Dataset Card for "Bundestagsreden_Paragraphs_with_likes"
More Information needed
ParagraphProof
ParagraphProof
tags: detection, paragraph, evidence-verification
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
This dataset contains paragraphs from various sources with the goal of identifying whether there is evidence supporting a specific claim within each paragraph. The dataset has been designed to aid in the development and training of machine learning models for the task of evidence-verification in textual data. Each… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphProof.ParagraphSentinel
ParagraphSentinel
tags: Claim Analysis, Text Parsing, Anomaly Detection
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'ParagraphSentinel' dataset is designed for the task of identifying and classifying claims within a text paragraph. Each entry in the dataset includes a paragraph of text and a label that indicates whether a claim is present and the type of claim identified. The dataset is suitable for machine learning models… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphSentinel.d1_science_long_paragraphs_1k_eval_636d
mlfoundations-dev/d1_science_long_paragraphs_1k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
18.7
54.5
77.0
29.2
37.8
36.7
20.2
4.1
5.5
AIME24
Average Accuracy: 18.67% ± 1.26%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
23.33%
7
30
3
23.33%
7
30
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_science_long_paragraphs_1k_eval_636d.
