datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs:
The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date.
We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.20_Newsgroups_Fixed
Dataset Card for 20_Newsgroups_Fixed
Dataset Summary
This dataset is a version of the 20 Newsgroups dataset fixed with the help of the Galileo ML Data Intelligence Platform. In a matter of minutes, Galileo enabled us to uncover and fix a multitude of errors within the original dataset. In the end, we present this improved dataset as a new standard for natural language experimentation and benchmarking using the Newsgroups dataset.
Curation Rationale
This… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/20_Newsgroups_Fixed.20_newsgroups
Dataset Card for "20_newsgroups"
More Information needed
20_newsgroups_Llama-3.1-8B-Instruct_vocab_2000_last20newsgroups_embedded
Dataset Card for 20-Newsgroups Embedded
This provides a subset of 20-Newsgroup posts, along with sentence embeddings, and a dimension reduced 2D data map.
This provides a basic setup for experimentation with various neural topic modelling approaches.
Dataset Details
Dataset Description
This is a dataset containing posts from the classic 20-Newsgroups dataset, along with sentence embeddings, and a dimension reduced 2D data map.
Per the source:
The… See the full description on the dataset page: https://huggingface.co/datasets/lmcinnes/20newsgroups_embedded.20_newsgroups20newsgroups20-Newsgroups
20 Newsgroups
Train
Some measurable characteristics of the dataset:
D — number of documents
W — modality dictionary size (number of unique tokens)
len D — average document length in modality tokens (number of tokens)
len D uniq — average document length in unique modality tokens (number of unique tokens)
D
@lemmatized W
@lemmatized len D
@lemmatized len D uniq
@bigram W
@bigram len D
@bigram len D uniq
value
11301
1.0614e+06
93.9204
60.5687
213701
18.9099… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/20-Newsgroups.20_newsgroups
Dataset Card for "20_newsgroups"
More Information needed
20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_2000_last20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_500_last20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_2000_last_variant_120_newsgroups_Llama-3.2-1B-Instruct_vocab_2000_last20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_1000_last20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_4000_last20newsgroups-paraphrased
20newsgroups-paraphrased
Paraphrased version of the 20 Newsgroups dataset. The body of each post has been paraphrased while the original email headers (From, Subject, Organization, etc.) are preserved verbatim. Topic labels are unchanged.
Model
Paraphrases were generated using Qwen/Qwen3-30B-A3B-Instruct-2507.
Original Dataset
Source: scikit-learn 20newsgroups
Task: Topic Classification
Classes: 20 (0-19 (20 newsgroup categories))
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/concretejungles/20newsgroups-paraphrased.20_newsgroups_demo
20newsgroups Demo
This 20 newgroups dataset is a filtered version containing only the "atheism.alt" and "soc.religion.christian". It is based on the SetFit/20_newsgroups dataset.
This dataset is used for our workshops at the AI Maker Community, a project sponsored by the Federal Ministry of Education and Research in Germany.
Original Dataset
This is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs:
The 20 newsgroups… See the full description on the dataset page: https://huggingface.co/datasets/aihpi/20_newsgroups_demo.20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_2000_last_variant_220_newsgroups_Qwen3.5-0.8B_vocab_2000_last20NewsGroups20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_2000_last_variant_320_newsgroups_lemma_test20newsgroups_topics
Dataset Card for 20-Newsgroups Embedded
This provides a subset of 20-Newsgroup posts, along with sentence embeddings, and a dimension reduced 2D data map.
This provides a basic setup for experimentation with various neural topic modelling approaches. Also provided is a set of topic
layers, giving topic names to clusters at five different layers of cluster resolution.
Dataset Details
Dataset Description
This is a dataset containing posts from the classic… See the full description on the dataset page: https://huggingface.co/datasets/lmcinnes/20newsgroups_topics.20_newsgroups_lemma_train20_newsgroups_triplet20_newsgroups_keywordscls_20newsgroups_SubjectTextVsLabel__BaseDefaultcls_20newsgroups_TextVsLabel__BaseDefault20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_2000_last_variant_420_newsgroups_Phi-3-mini-128k-instruct_vocab_2000_last
