datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
forecast-news
Forecast News
Deduplicated daily news corpus used by forecast-sim and future-sim.
Snapshot
31,859,020 articles
3,463 daily partitions
Coverage: 2016-08-26 through 2026-08-31
Snapshot published: 2026-09-18
Stored data size: approximately 158.5 GiB
The Parquet files are the canonical complete representation. The repository also
contains daily JSONL files where available and compact headline JSON files used
by article-browsing workflows.
Layout
Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.telegram-news-ua-dataset
Aisberg Telegram News UA
A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs:
The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date.
We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.ag_newsnewswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.bbc-news
BBC News Topic Dataset
Dataset on BBC News Topic Classification consisting of 2,225 articles published on the BBC News website corresponding during 2004-2005. Each article is labeled under one of 5 categories: business, entertainment, politics, sport or tech.
Original source for this dataset:
Derek Greene, Pádraig Cunningham, “Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering,” in Proc. 23rd International Conference on Machine learning (ICML’06)… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/bbc-news.IndustryCorpus_news[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_news.ag_newsAG's News Topic Classification Dataset
Version 3, Updated 09/09/2015
ORIGIN
AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search… See the full description on the dataset page: https://huggingface.co/datasets/sh0416/ag_news.news-category-datasetDataset from https://www.kaggle.com/datasets/rmisra/news-category-dataset
daily-bio-newsCSL-News
Summary
This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale".
CSL-News is a large-scale Chinese Sign Language dataset designed for developing robust sign language understanding models.
Code: https://github.com/ZechengLi19/Uni-Sign
Download
Please refer to download script to download CSL_News.
You can also download each file by wget, for instance:
wget… See the full description on the dataset page: https://huggingface.co/datasets/ZechengLi19/CSL-News.news-entertainment-datasetNews_Category_Dataset_v2news-politics-datasetnews-education-datasetnews-tech-datasetnews-finance-datasetnews21-instructionsSP500-Financial-News-Articles-Time-SeriesTextual Time Series Dataset for finetuning / pretraining.
Json version of original dataset.
Original Dataset : https://www.kaggle.com/datasets/skywalker290/financial-news-article-and-stock-trend-dataset?select=stock_data_articles.csv
news
News
Description
We scrape the news sites that publish content under CC BY or CC BY-SA according to opennewswire.
These include 360info, Africa is a Country,
Alt News,
Balkan Diskurs,
Factly,
Freedom of the Press Foundation,
Agenzia Fides,
Global Voices,
Meduza,
Mekong Eye,
Milwaukee Neighborhood News Service,
Minority Africa,
New Canadian Media,
SciDev.Net,
The Solutions Journalism Exchange,
Tasnim News Agency… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/news.CSL-News
Summary
This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale".
CSL-News is a large-scale Chinese Sign Language dataset designed for developing robust sign language understanding models.
Code: https://github.com/ZechengLi19/Uni-Sign
Download
Please refer to download script to download CSL_News.
You can also download each file by wget, for instance:
wget… See the full description on the dataset page: https://huggingface.co/datasets/Superinn/CSL-News.newspaper_navigatorgdelt-news-headlinesbmw-pressclub-news
BMW PressClub News Dataset
This dataset contains press releases and news articles scraped from
BMW PressClub.
Dataset Structure
JSON Format (bmw_articles.json)
{
"scraped_at": "2025-12-17T10:00:00",
"source": "https://www.press.bmwgroup.com/global/article",
"count": 100,
"articles": [
{
"title": "BMW presents the new X5",
"date": "17.12.2025",
"article_type": "Press Release",
"summary": "...",
"tags": ["BMW X5", "SUV"]… See the full description on the dataset page: https://huggingface.co/datasets/Alwin-Yang/bmw-pressclub-news.Sarcasm_News_HeadlinePast studies in Sarcasm Detection mostly make use of Twitter datasets collected using hashtag based supervision but such datasets are noisy in terms of labels and language. Furthermore, many tweets are replies to other tweets and detecting sarcasm in these requires the availability of contextual tweets.
To overcome the limitations related to noise in Twitter datasets, this Headlines dataset for Sarcasm Detection is collected from two news website. TheOnion aims at producing sarcastic versions… See the full description on the dataset page: https://huggingface.co/datasets/raquiba/Sarcasm_News_Headline.binhvq_news21_rawCSL-News_pose
Summary
This is the pose format dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale".
CSL-News is a large-scale Chinese Sign Language dataset designed for developing robust sign language understanding models.
Code: https://github.com/ZechengLi19/Uni-Sign
Download
Please refer to download script to download CSL_News.
You can also download each file by wget, for instance:
wget… See the full description on the dataset page: https://huggingface.co/datasets/ZechengLi19/CSL-News_pose.trec-newsrussian-news-telegram-dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic News and Media,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the channel. views… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian-news-telegram-dataset.newsquadfr
Dataset Card for newsquadfr
Dataset Summary
newsquadfr is a small dataset created for Question Answering task. Contexts are paragraphs of articles extracted from nine online french newspaper during year 2020/2021. newsquadfr stands for Newspaper question answering dataset in french. inspired by Piaf and Squad dataset. 2 520 triplets context - question - answer.
from datasets import load_dataset
ds_name = 'lincoln/newsquadfr'
# exemple 1
ds_newsquad =… See the full description on the dataset page: https://huggingface.co/datasets/lincoln/newsquadfr.
