datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.financial-news-articles-filtereddataset_info:
features:
- name: title
dtype: string
- name: text
dtype: string
- name: url
dtype: string
- name: word_count
dtype: int64
splits:
- name: train
num_bytes: 554834105.9892601
num_examples: 199711
download_size: 459025008
dataset_size: 554834105.9892601
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
brazilian-news-articlesmoi-news-articles-dataset
MOI News & Article Dataset 🇲🇲
This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.Ateso_news_articles
Ateso News Articles
Ateso (teo) is one of the most spoken languages in Uganda
Dataset Details
Artictles were scrapped from https://www.aicerit.co.ug
Luganda_news_articles
Luganda News Articles
Luganda (lug) is one of the most spoken languages in Uganda.
Scrapped from https://www.bukedde.co.ug/ & https://gambuuze.ug/
BBC_Eng_News_Articles_dataset
BBC News Articles Dataset
Dataset Description
A collection of 2,225 news articles from BBC, suitable for text classification, summarization, and NLP tasks.
Dataset Summary
Metric
Value
Total Articles
2,225
Unique Articles
2,092
Columns
filename, article_text
Language
English
Source
BBC News
Dataset Structure
Data Fields
Field
Type
Description
filename
string
Unique identifier/filename for each… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/BBC_Eng_News_Articles_dataset.
