datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs:
The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date.
We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.newsgroupThe 20 Newsgroups data set is a collection of approximately 20,000 newsgroup documents, partitioned (nearly) evenly across
20 different newsgroups. The 20 newsgroups collection has become a popular data set for experiments in text applications of
machine learning techniques, such as text classification and text clustering.20_Newsgroups_Fixed
Dataset Card for 20_Newsgroups_Fixed
Dataset Summary
This dataset is a version of the 20 Newsgroups dataset fixed with the help of the Galileo ML Data Intelligence Platform. In a matter of minutes, Galileo enabled us to uncover and fix a multitude of errors within the original dataset. In the end, we present this improved dataset as a new standard for natural language experimentation and benchmarking using the Newsgroups dataset.
Curation Rationale
This… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/20_Newsgroups_Fixed.20_newsgroups
Dataset Card for "20_newsgroups"
More Information needed
20_newsgroups_Llama-3.1-8B-Instruct_vocab_2000_last20_newsgroups20-Newsgroups
20 Newsgroups
Train
Some measurable characteristics of the dataset:
D — number of documents
W — modality dictionary size (number of unique tokens)
len D — average document length in modality tokens (number of tokens)
len D uniq — average document length in unique modality tokens (number of unique tokens)
D
@lemmatized W
@lemmatized len D
@lemmatized len D uniq
@bigram W
@bigram len D
@bigram len D uniq
value
11301
1.0614e+06
93.9204
60.5687
213701
18.9099… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/20-Newsgroups.20_newsgroups
Dataset Card for "20_newsgroups"
More Information needed
newsgroups
Dataset Card for "20-Newsgroups"
20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_2000_lastprompted-newsgroupsfewshot-prompted-newsgroups20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_500_last20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_2000_last_variant_1newsgroups20_newsgroups_Llama-3.2-1B-Instruct_vocab_2000_last20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_1000_last20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_4000_lastnewsgroups-mininewsgroups-miniThe data in this dataset is a subset of 20newsgroups/SciKit dataset:
https://scikit-learn.org/0.19/modules/generated/sklearn.datasets.fetch_20newsgroups.html#sklearn.datasets.fetch_20newsgroups
license: mit
dataset_info:
pretty_name: 'SciKit newsgroup20 subset'
features:
- name: index
dtype: int64
- name: Text
dtype: string
- name: Label
dtype: int32
- name: Class Name
dtype: string
task_categories:
-text classification
-sentence similarity
tags:… See the full description on the dataset page: https://huggingface.co/datasets/acloudfan/newsgroups-mini.20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_2000_last_variant_2llm-eval-twenty_newsgroups_v2newsgroups20_newsgroups_demo
20newsgroups Demo
This 20 newgroups dataset is a filtered version containing only the "atheism.alt" and "soc.religion.christian". It is based on the SetFit/20_newsgroups dataset.
This dataset is used for our workshops at the AI Maker Community, a project sponsored by the Federal Ministry of Education and Research in Germany.
Original Dataset
This is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs:
The 20 newsgroups… See the full description on the dataset page: https://huggingface.co/datasets/aihpi/20_newsgroups_demo.newsgroups7_balanced20_newsgroups_Qwen3.5-0.8B_vocab_2000_lastnewsgroups20_newsgroups_ERNIE-4.5-0.3B-PT_vocab_2000_last_variant_320_newsgroups_lemma_testnewsgroups_Full-p_1
