datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
news-political-bias-classification-datasetDataset actually from kaggle.
Couldn't find it here so I uploaded it.
BilCat-news-classificationBilCat: Bilkent Text Classification (News Categorization) Dataset
7540 Turkish news articles (Milliyet and TRT merged) with category labels (Dunya, Ekonomi, Politika, KulturSanat, Saglik, Spor, Turkiye, Yazarlar).
Column header is the first line.
Other details are at https://github.com/BilkentInformationRetrievalGroup/BilCat/
Citation:
C. Toraman, F. Can and S. Koçberber. Developing a text categorization template for Turkish news portals. 2011 International Symposium on Innovations in… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/BilCat-news-classification.Amharic-News-Text-classification-Dataset
An Amharic News Text classification Dataset
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.malayalam_news_classificationSinhala-News-Category-classificationThis file contains news texts (sentences) belonging to 5 different news categories (political, business, technology, sports and Entertainment). The original dataset was released by Nisansa de Silva (Sinhala Text Classification: Observations from the Perspective of a Resource Poor Language, 2015). The original dataset is processed and cleaned of single word texts, English only sentences etc.
If you use this dataset, please cite {Nisansa de Silva, Sinhala Text Classification: Observations from… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Category-classification.gujarati_news_classificationtamil_news_classificationSinhala-News-Source-classificationThis dataset contains Sinhala news headlines extracted from 9 news sources (websites) (Sri Lanka Army, Dinamina, GossipLanka, Hiru, ITN, Lankapuwath, NewsLK,
Newsfirst, World Socialist Web Site-Sinhala). This is a processed version of the corpus created by Sachintha, D., Piyarathna, L., Rajitha, C., and Ranathunga, S. (2021). Exploiting parallel corpora to improve multilingual embedding based document and sentence alignment. Single word sentences, invalid characters have been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Source-classification.Fake-News-ClassificationDevelop a machine learning program to identify when an article might be fake news. Run by the UTK Machine Learning Club.
This is the Dataset to the Fake-News-Classifier competition in Kaggle. There is a Test csv to check for predictions.
Citation
William Lifferth. (2018). Fake News. Kaggle. https://kaggle.com/competitions/fake-news
abusive-news-comment-ind-classificationref: https://github.com/dhamirdesrul/Indonesian-Online-News-Comments
punjabi_news_classificationodia_news_classificationnews-tam-classificationref: https://www.kaggle.com/datasets/disisbig/bengali-news-dataset
murasu-news-tam-classificationref: https://www.kaggle.com/datasets/vijayabhaskar96/tamil-news-classification-dataset-tamilmurasu
news-khm-classificationref: https://github.com/phylypo/khmer-text-data
kannada_news_classificationNews-Classification-Datasetmarathi_news_classificationag-news-topic-classification-processed
AG News Topic Classification Processed Dataset
This dataset is a processed subset of the AG News Classification Dataset from Kaggle.
Task
The task is news topic classification. Each example contains a news title and description combined into a single text field. The goal is to classify each news text into one of four categories:
World
Sports
Business
Sci/Tech
Data Source
The original data comes from the AG News Classification Dataset on Kaggle. For this… See the full description on the dataset page: https://huggingface.co/datasets/Thisisruiii/ag-news-topic-classification-processed.telugu_news_classificationnbc_headlines.csvtelugu_news_classificationnews-classificationThis is edit dataset from
https://www.kaggle.com/datasets/banuprakashv/news-articles-classification-dataset-for-nlp-and-ml/data
Have edited to fine tune LLM's
zero-shot-classification-news-uzbeknews_classification_expanded10,000 total news headlines.
5000 fox. 5000 nbc.
Amharic-News-Classificationfake-news-classification-datasetheadline_datatrain_data_nbc_skewednews-classification-labeled
