datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
political_bias_in_news_articlesfinancial-news-articles
Dataset Card for "financial-news-articles"
More Information needed
The data was obtained from here
SP500-Financial-News-Articles-Time-SeriesTextual Time Series Dataset for finetuning / pretraining.
Json version of original dataset.
Original Dataset : https://www.kaggle.com/datasets/skywalker290/financial-news-article-and-stock-trend-dataset?select=stock_data_articles.csv
news-articles-ptbr-dataset
Dataset Card for "news-articles-ptbr-dataset"
More Information needed
aya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.ai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.financial-news-articles-filtereddataset_info:
features:
- name: title
dtype: string
- name: text
dtype: string
- name: url
dtype: string
- name: word_count
dtype: int64
splits:
- name: train
num_bytes: 554834105.9892601
num_examples: 199711
download_size: 459025008
dataset_size: 554834105.9892601
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
climate-news-articles
🌍 Jeu de données d'articles de presse française labellisés comme traitant ou non des sujets liés au climat
🇬🇧 / 🇺🇸 : as this data set is based only on French data, all explanations are written in French in this repository. The goal of the dataset is to train a model to classify titles of French newspapers in two categories : if it's about climate or not.
🗺️ Le contexte
Ce jeu de données de classification de titres d'article de presse française a été réalisé pour… See the full description on the dataset page: https://huggingface.co/datasets/pierre-loic/climate-news-articles.mteb-nl-news-articles-clsThis dataset contains Dutch news articles along with their corresponding categories, sourced from the Nederlandse Oproep Stichting.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch},
author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-cls.all-the-news-2-pythia-tfidf-topic-stratified-v1-articlesall-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articleschatgpt-news-articles
Dataset Card for "chatgpt-news-articles"
Dataset Summary
The ChatGPT CNN / DailyMail Dataset is a small sample of the original CNN / DailyMaily English-language dataset containing 25k unique news articles. For each corresponding article written by journalists at CNN and the Daily Mail, there is an article written by ChatGPT using the highlights provided by human annotators. The current version supports can be used to study the language comparison between human and ChatGPT… See the full description on the dataset page: https://huggingface.co/datasets/isarth/chatgpt-news-articles.ai-jobs-news-articles-abstracts
News articles and research abstracts on AI, labor, and jobs
Dataset summary
This file is a standalone CSV of news articles (full scraped text) and scholarly paper abstracts curated for research on artificial intelligence, work, and labor markets. Each row is one document: a stable id, publication date, normalized title and main text, and a small metadata dictionary.
Rows: 53,526
document_class
Rows
Approx. date range (date column)
news
29,857
Jan. 2025… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/ai-jobs-news-articles-abstracts.News_Articles_Categorization
Dataset Card for News_Articles_Categorization
Dataset Description
3722 News Articles classified into different categories namely: World, Politics, Tech, Entertainment, Sport, Business, Health, and Science
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely Text and Category.
The Text column consists of the news article and the Category column consists of the class each article belongs to… See the full description on the dataset page: https://huggingface.co/datasets/valurank/News_Articles_Categorization.CNN_News_Articles_2011-2022
CNN News Articles 2011-2022 Dataset
Introduction
This dataset contains CNN News Articles from 2011 to 2022 after basic cleaning. The dataset includes the following information:
Category
Full text
The data was downloaded from Kaggle at this URL: https://www.kaggle.com/datasets/hadasu92/cnn-articles-after-basic-cleaning. The dataset was split into two sets:
Train set with 32,218 examples
Test set with 5,686 examples
Usage
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/CNN_News_Articles_2011-2022.news_articles
Dataset Card for "news_articles"
More Information needed
reddit_news_articles_commentsHinduTamil-News-Articles-Dataset
HinduTamil News Articles Dataset
Overview
This dataset contains news articles in Tamil language scraped from the Hindu Tamil news website. Each article includes its title, author, city, published date, and text.
Motivation
This dataset was created to provide a comprehensive collection of Tamil news articles for research and analysis purposes.
Data Sources and collection method
The data in this dataset was collected from the Hindu Tamil news website… See the full description on the dataset page: https://huggingface.co/datasets/Shwetasss/HinduTamil-News-Articles-Dataset.shipping_news_articles_summary_embbcms-fake-news-articlesmteb-nl-news-articles-retThis dataset contains Dutch news articles, sourced from the Nederlandse Oproep Stichting.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch},
author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-ret.bloomberg-news-articles-pretraining-datasetturkish-news-articles-sequence-classification
Turkish Sentence Continuation Dataset
Dataset Summary
This dataset is a binary sentence-pair classification dataset created from Turkish newspaper articles.
Each example consists of two sentences:
Positive (label = 1): the second sentence directly follows the first sentence in the original article.
Negative (label = 0): the second sentence is unrelated and comes from a different context.
The dataset is suitable for sentence coherence, discourse understanding, and… See the full description on the dataset page: https://huggingface.co/datasets/oguzinc/turkish-news-articles-sequence-classification.shipping_news_articles_lsathree_line_summarization_for_japanese_news_articlesライブドアニュースコーパスの3行要約データセットです。
Llama v2向けのプロンプトを追加して成形してあります。
学習に利用する際は、 [R_START] [R_END] をspecial tokenとして追加することを推奨します。
Number of rows: 3,907
Datasetは以下のリポジトリを利用してscrapeしました。
git@github.com:KodairaTomonori/ThreeLineSummaryDataset.git
bloomberg-news-articles-pretraining-datasetnews_articles_daily_mall-the-news-2-tfidf-invfreq-topic-stratified-v1-articlesnews_articles
News Articles Classification Dataset
This dataset consists of news articles labeled with corresponding categories for classification tasks.
Overview
The news articles classification dataset is a collection of articles sourced from various news outlets, each labeled with a specific category. The dataset is designed for tasks such as text classification, topic modeling, and sentiment analysis.
Dataset Information
Name: News Articles Classification Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bushra1dajam/news_articles.shipping_news_articles
