datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ILUR-news-text-classification-corpus
News Texts Dataset
We release a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society. The articles are split into train (2242k tokens) and test sets (425k tokens).
For more details, refer to the paper.
news-political-bias-classification-datasetDataset actually from kaggle.
Couldn't find it here so I uploaded it.
thura-myanmar-news-mya-classificationref: https://huggingface.co/datasets/ThuraAung1601/myanmar_news
news-lao-classification
News_lao_Classification
Deduplicated copy of kornwtp/news-lao-classification.
Splits
split
rows
test
3,062
train
9,161
validation
3,062
BilCat-news-classificationBilCat: Bilkent Text Classification (News Categorization) Dataset
7540 Turkish news articles (Milliyet and TRT merged) with category labels (Dunya, Ekonomi, Politika, KulturSanat, Saglik, Spor, Turkiye, Yazarlar).
Column header is the first line.
Other details are at https://github.com/BilkentInformationRetrievalGroup/BilCat/
Citation:
C. Toraman, F. Can and S. Koçberber. Developing a text categorization template for Turkish news portals. 2011 International Symposium on Innovations in… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/BilCat-news-classification.murasu-news-tam-classification
MurasuNews_tam_Classification
Deduplicated copy of kornwtp/murasu-news-tam-classification.
Splits
split
rows
train
126,675
Amharic-News-Text-classification-Dataset
An Amharic News Text classification Dataset
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.malayalam_news_classificationnews-khm-classification
News_khm_Classification
Deduplicated copy of kornwtp/news-khm-classification.
Splits
split
rows
train
818
news-tam-classification
News_tam_Classification
Deduplicated copy of kornwtp/news-tam-classification.
Splits
split
rows
train
5,309
validation
1,333
BLUGE-bengali-news-classification
BLUGE-NCC: Bangla News Classification
BLUGE-NCC is a meticulously curated and balanced Bangla News Category Classification dataset, one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it.
Dataset Description
This task classifies Bangla news articles… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BLUGE-bengali-news-classification.abusive-news-comment-ind-classification
AbusiveNewsComment_ind_Classification
Deduplicated copy of kornwtp/abusive-news-comment-ind-classification.
Splits
split
rows
train
3,164
news-sentiment-zsm-classification
NewsSentiment_zsm_Classification
Deduplicated copy of kornwtp/news-sentiment-zsm-classification.
Splits
split
rows
train
3,673
synthetic-text-classification-news
Dataset Card for synthetic-text-classification-news
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/argilla/synthetic-text-classification-news/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news.synthetic-text-classification-news-multi-label
Dataset Card for synthetic-text-classification-news-multi-label
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/synthetic-text-classification-news-multi-label/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news-multi-label.news-mya-classification
News_mya_Classification
Deduplicated copy of kornwtp/news-mya-classification.
Splits
split
rows
train
2,042
text-classification-news-topics
Dataset Card for test
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/test/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/text-classification-news-topics.Sinhala-News-Category-classificationThis file contains news texts (sentences) belonging to 5 different news categories (political, business, technology, sports and Entertainment). The original dataset was released by Nisansa de Silva (Sinhala Text Classification: Observations from the Perspective of a Resource Poor Language, 2015). The original dataset is processed and cleaned of single word texts, English only sentences etc.
If you use this dataset, please cite {Nisansa de Silva, Sinhala Text Classification: Observations from… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Category-classification.amharic-news-category-classification
Amharic News Category Classification
This amharic text dataset can be used to train/finetune models for the following tasks
classification : using the categories
summarization : using the headlines
Finetuning
Here is a github repo that contains three notebooks that use this dataset to finetune the following models.
xlm-roberta-base : a multilingual transformer model with 280M parameters
bert-small-amharic : a new amharic version of the bert-small transformer model… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-news-category-classification.news-mya-classificationref: https://huggingface.co/datasets/mteb/MyanmarNews
thura-myanmar-news-mya-classification
ThuraMyanmarNews_mya_Classification
Deduplicated copy of kornwtp/thura-myanmar-news-mya-classification.
Splits
split
rows
test
7,238
train
7,238
vietnamese-news-classification
Vietnamese News Classification Dataset (1.3M)
Dataset Description
This dataset contains approximately 1.3 million Vietnamese news articles collected from major online news portals. Structured similarly to the popular AG News dataset, it serves as a valuable resource for experimenting with multi-class text classification in Vietnamese.
The dataset covers 11 topics (categories) ranging from current affairs, sports, technology, to entertainment.
Curated by: Nam Syntax… See the full description on the dataset page: https://huggingface.co/datasets/NamSyntax/vietnamese-news-classification.news-classificationgujarati_news_classificationtamil_news_classificationen_si_news_classificationKhmer_News_classificationSinhala-News-Source-classificationThis dataset contains Sinhala news headlines extracted from 9 news sources (websites) (Sri Lanka Army, Dinamina, GossipLanka, Hiru, ITN, Lankapuwath, NewsLK,
Newsfirst, World Socialist Web Site-Sinhala). This is a processed version of the corpus created by Sachintha, D., Piyarathna, L., Rajitha, C., and Ranathunga, S. (2021). Exploiting parallel corpora to improve multilingual embedding based document and sentence alignment. Single word sentences, invalid characters have been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Source-classification.news-sentiment-zsm-classificationref: https://github.com/mesolitica/malaysian-dataset/tree/master/sentiment/news-sentiment
turkish-news-articles-sequence-classification
Turkish Sentence Continuation Dataset
Dataset Summary
This dataset is a binary sentence-pair classification dataset created from Turkish newspaper articles.
Each example consists of two sentences:
Positive (label = 1): the second sentence directly follows the first sentence in the original article.
Negative (label = 0): the second sentence is unrelated and comes from a different context.
The dataset is suitable for sentence coherence, discourse understanding, and… See the full description on the dataset page: https://huggingface.co/datasets/oguzinc/turkish-news-articles-sequence-classification.
