datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
textclassificationMNLItext_classificationend2end_textclassification
Dataset Card for end2end_textclassification
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/argilla/end2end_textclassification.end2end_textclassification_with_suggestions_and_responses
Dataset Card for end2end_textclassification_with_suggestions_and_responses
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure… See the full description on the dataset page: https://huggingface.co/datasets/argilla/end2end_textclassification_with_suggestions_and_responses.text-classification-subject
Dataset Card for "text-classification-subject"
More Information needed
text-classification-news-topics
Dataset Card for test
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/test/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/text-classification-news-topics.text-classification-dataset-exampleText-Classification-and-Relation-Event-Extraction-Mix-datasetsThe paper of GIELLM dataset.
https://arxiv.org/abs/2311.06838
Cite:
@article{gan2023giellm,
title={Giellm: Japanese general information extraction large language model utilizing mutual reinforcement effect},
author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori},
journal={arXiv preprint arXiv:2311.06838},
year={2023}
}
The dataset constructed base in livedoor news corpus 関口宏司 https://www.rondhuit.com/download.html
end2end_textclassification_with_vectors
Dataset Card for end2end_textclassification_with_vectors
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when… See the full description on the dataset page: https://huggingface.co/datasets/argilla/end2end_textclassification_with_vectors.text_classificationend2end_textclassification
Dataset Card for end2end_textclassification
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/carlosug/end2end_textclassification.end2end_textclassification_with_metadata
Dataset Card for end2end_textclassification_with_metadata
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when… See the full description on the dataset page: https://huggingface.co/datasets/argilla/end2end_textclassification_with_metadata.text-classification-comparison
Text Classification Comparison: Supervised vs Unsupervised on stanfordnlp/imdb
Dataset
stanfordnlp/imdb (Maas et al., 2011)
50,000 IMDB movie reviews: 25,000 train / 25,000 test
Binary sentiment: 0 (negative) / 1 (positive), perfectly balanced in both splits
Preprocessing: TF-IDF (15,000 features, bigrams, sublinear TF, min_df=3)
Models Chosen
Model
Type
Key Reference
Why
Logistic Regression
Linear supervised
McFadden (1974); Ng &… See the full description on the dataset page: https://huggingface.co/datasets/IntimateUser6969/text-classification-comparison.text-classificationtext-classification-checkpoint-downloadsText_Classification_Deutsch_Beispieltext_classification2000_TextClassification
Dataset Card for "2000_TextClassification"
More Information needed
end2end_textclassification
Dataset Card for "end2end_textclassification"
More Information needed
Text_classification_by_subject_area
🇰🇿 Kazakh Topic and Domain Identification Dataset
Dataset Summary
Kazakh Topic and Domain Identification Dataset is a Kazakh-language instruction-following dataset designed for topic recognition, domain classification, and text understanding tasks.
Each sample contains a short Kazakh prompt, a long Kazakh text passage, a target response, a domain label, and a unique sample identifier. The dataset is intended to help Large Language Models (LLMs) and NLP systems… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Text_classification_by_subject_area.text-classification-chatGPT-100text_classification_testtext_classification_dataset1000_TextClassification
Dataset Card for "1000_TextClassification"
More Information needed
text-classification-pl
Labels
0: normal
1: toxic
Text_Classification_DatasetText-classification-260ktraining-textclassificationtext-classification-datasettext_classification
