CoolFace
Datasetpublic

concretejungles/20newsgroups-paraphrased

20newsgroups-paraphrased Paraphrased version of the 20 Newsgroups dataset. The body of each post has been paraphrased while the original email headers (From, Subject, Organization, etc.) are preserved verbatim. Topic labels are unchanged. Model Paraphrases were generated using Qwen/Qwen3-30B-A3B-Instruct-2507. Original Dataset Source: scikit-learn 20newsgroups Task: Topic Classification Classes: 20 (0-19 (20 newsgroup categories))… See the full description on the dataset page: https://huggingface.co/datasets/concretejungles/20newsgroups-paraphrased.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes31downloads
Dataset Card

20newsgroups-paraphrased

Paraphrased version of the 20 Newsgroups dataset. The body of each post has been paraphrased while the original email headers (From, Subject, Organization, etc.) are preserved verbatim. Topic labels are unchanged.

Model

Paraphrases were generated using [Qwen/Qwen3-30B-A3B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-30B-A3B-Instruct-2507).

Original Dataset

Dataset Structure

SplitExamples
train9,051
validation2,263
test7,532

Columns

  • —text (string): The paraphrased text.
  • —label (int): The original label, preserved from the source dataset.
  • —original_text (string): The original text before paraphrasing.

How to Use

python
from datasets import load_dataset

ds = load_dataset("concretejungles/20newsgroups-paraphrased")
print(ds["train"][0])

Generation Details

  • —Paraphrase model: Qwen/Qwen3-30B-A3B-Instruct-2507
  • —Method: Each example was paraphrased with a dataset-specific prompt designed to preserve the label semantics.
  • —Note: Email headers (From, Subject, Organization, etc.) are kept verbatim; only the post body is paraphrased.