concretejungles/20newsgroups-paraphrased
20newsgroups-paraphrased Paraphrased version of the 20 Newsgroups dataset. The body of each post has been paraphrased while the original email headers (From, Subject, Organization, etc.) are preserved verbatim. Topic labels are unchanged. Model Paraphrases were generated using Qwen/Qwen3-30B-A3B-Instruct-2507. Original Dataset Source: scikit-learn 20newsgroups Task: Topic Classification Classes: 20 (0-19 (20 newsgroup categories))… See the full description on the dataset page: https://huggingface.co/datasets/concretejungles/20newsgroups-paraphrased.
20newsgroups-paraphrased
Paraphrased version of the 20 Newsgroups dataset. The body of each post has been paraphrased while the original email headers (From, Subject, Organization, etc.) are preserved verbatim. Topic labels are unchanged.
Model
Paraphrases were generated using [Qwen/Qwen3-30B-A3B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-30B-A3B-Instruct-2507).
Original Dataset
- Source: scikit-learn 20newsgroups
- Task: Topic Classification
- Classes: 20 (0-19 (20 newsgroup categories))
Dataset Structure
Columns
text(string): The paraphrased text.label(int): The original label, preserved from the source dataset.original_text(string): The original text before paraphrasing.
How to Use
from datasets import load_dataset
ds = load_dataset("concretejungles/20newsgroups-paraphrased")
print(ds["train"][0])Generation Details
- Paraphrase model:
Qwen/Qwen3-30B-A3B-Instruct-2507 - Method: Each example was paraphrased with a dataset-specific prompt designed to preserve the label semantics.
- Note: Email headers (From, Subject, Organization, etc.) are kept verbatim; only the post body is paraphrased.
