datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emakhuwa-News-Topic-ClassificationBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-News-Topic-Classification.vi-news-4topics-classification
Vietnamese News Classification Dataset (4 Topics)
Description
This dataset contains Vietnamese news headlines collected from publicly available news sources. The dataset is designed for text classification tasks and multilingual NLP research.
Task
Multi-class text classification (4 topics): world, sport, business, tech
Data Collection
The data was collected from Vietnamese news websites (e.g., VnExpress). Only headlines were used to ensure… See the full description on the dataset page: https://huggingface.co/datasets/ngocleltt/vi-news-4topics-classification.FT_news_classification
