datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-paragraphs
Wikipedia Paragraph Samples
Dataset Description
This dataset contains paragraphs extracted from randomly selected English Wikipedia articles. It provides a diverse sample of Wikipedia content across various topics.
Dataset Details
Name: Wikipedia Paragraph Samples
Version: 1.0
Date Created: 2024-08-20
Language: English
Format: JSONLines
Contents
Each line in the dataset represents a single paragraph and contains two fields:
Title of the Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs.wikipedia-first-paragraphwiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.wiki_paragraphs_english
WIKI Paragraphs English
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.Synthetic-Pretrain-Paragraphs-150Topics
Synthetic-Pretrain-Paragraphs-150Topics
A synthetic dataset consisting of continuous text paragraphs in Russian and English, generated using Qwen2.5-7B-Instruct.
Dataset Curation
Source: Generated via vLLM on an RTX 3060.
Topics: Covers 150 fundamental fields of knowledge, ranging from programming to speleology.
Constraint: Strict prohibition on greetings, lists, and conversational fillers; contains only raw factual descriptions.
Warning: As pure synthetic data, it… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/Synthetic-Pretrain-Paragraphs-150Topics.wikipedia-paragraphs-complete
Wikipedia Paragraphs Complete Dataset
This dataset consists of English Wikipedia paragraphs ranging from 1 000 to 8 000 characters in length. It was sourced from the Wikimedia dump: "wikimedia/wikipedia", "20231101.en".
Preprocessing Steps
The dataset has undergone extensive cleaning and normalization, including:
Removing brackets
Removing HTML tags
Normalizing bullet points, hyphenated words, quotation marks, Unicode characters, and whitespace
Replacing email… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-complete.wikipedia-paragraph-summaries
Wikipedia Paragraph Summaries Dataset
The Wikipedia Paragraph Summaries Dataset is designed for the task of text summarization, specifically generating concise summaries from paragraphs extracted from English Wikipedia articles. Each entry in the dataset consists of an input paragraph and its corresponding summary, facilitating research in natural language processing (NLP) and machine learning.
Data Format:
The dataset is provided in JSON Lines format (.jsonl), where each line… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraph-summaries.dfm12-reordering-integrated-nl-paragraph-reordering
dfm12-reordering-integrated-nl-paragraph-reordering
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-reordering-integrated-nl-paragraph-reordering.dfm12-reordering-integrated-nn-paragraph-reordering
dfm12-reordering-integrated-nn-paragraph-reordering
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-reordering-integrated-nn-paragraph-reordering.dfm12-reordering-integrated-nb-paragraph-reordering
dfm12-reordering-integrated-nb-paragraph-reordering
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-reordering-integrated-nb-paragraph-reordering.wikipedia-paragraph-sft
Wikipedia Paragraph Supervised Finetuning Dataset
Model Description
This dataset is designed for training language models to generate supervised finetuning data from raw text. It consists of text passages and corresponding question-answer pairs in JSONLines format.
Intended Use
The primary purpose of this dataset is to enable large language models (LLMs) to generate high-quality supervised finetuning data from raw text inputs, useful for creating custom… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraph-sft.dfm12-reordering-integrated-sv-paragraph-reordering
dfm12-reordering-integrated-sv-paragraph-reordering
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-reordering-integrated-sv-paragraph-reordering.wikipedia-paragraphs-direct-paraphrases
Wikipedia Paragraphs Direct Paraphrases
Paraphrases of Wikipedia paragraphs using AI large language models.
Paragraphs from agentlans/wikipedia-paragraphs-complete sample_k10000 and sample_k50000 splits
Paraphrased using Qwen/Qwen3.5-9B and a distilled Qwen/Qwen3-4B-Instruct-2507 with the following prompt:
Rewrite the following paragraph entirely in your own words while preserving every fact, detail, meaning, nuance, and level of specificity. Do not add, remove, reinterpret… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-direct-paraphrases.task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph.bangla-word-to-paragraphwikitext_with_entitled_paragraphswikipedia-paragraph-conversation
Wikipedia Paragraphs Conversation Dataset
This dataset contains approximately 800 high-quality English Wikipedia paragraphs paired with synthetic, multi-turn conversations. It is specifically designed to train or fine-tune LLMs to act as synthetic data generators—transforming static knowledge into natural, multi-turn dialogue.
Dataset Summary
The goal of this dataset is to bridge the gap between raw factual prose and interactive conversational formats. Each row features a… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraph-conversation.michael_paragraphs_units_5
michael_paragraphs_units_5
Dataset uploaded with Python via huggingface_hub.
Files
JSONL source file uploaded to this repository
Notes
Custom dataset
Uploaded automatically from a local file
