datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
neweag_news
Dataset Card for "ag_news"
Dataset Summary
AG is a collection of more than 1 million news articles. News articles have been
gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of
activity. ComeToMyHead is an academic news search engine which has been running
since July, 2004. The dataset is provided by the academic comunity for research
purposes in data mining (clustering, classification, etc), information retrieval
(ranking, search, etc)… See the full description on the dataset page: https://huggingface.co/datasets/fancyzhx/ag_news.infini-news-corpus
INFINI-NEWS Corpus
🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.forecast-news
Forecast News
Deduplicated daily news corpus used by forecast-sim and future-sim.
Snapshot
31,859,020 articles
3,463 daily partitions
Coverage: 2016-08-26 through 2026-08-31
Snapshot published: 2026-09-18
Stored data size: approximately 158.5 GiB
The Parquet files are the canonical complete representation. The repository also
contains daily JSONL files where available and compact headline JSON files used
by article-browsing workflows.
Layout
Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.newyorker_caption_contest
Dataset Card for New Yorker Caption Contest Benchmarks
Dataset Summary
See capcon.dev for more!
Data from:
Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
@inproceedings{hessel2023androids,
title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding''
Benchmarks from {The New Yorker Caption Contest}},
author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian
and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.new-york-smells
New York Smells: A Large Multimodal Dataset for Olfaction
While olfaction is central to how animals perceive the world, this rich chemical
sensory modality remains largely inaccessible to machines. One key bottleneck is the
lack of diverse, multimodal olfactory data collected in natural settings. We present
New York Smells, a large-scale dataset of paired image and olfactory signals
captured in-the-wild. Our dataset contains 7,000 smell-image pairs from 3,500 distinct
objects… See the full description on the dataset page: https://huggingface.co/datasets/cvlab/new-york-smells.dojo_stock_news
Languages: 简体中文 · English
dojo_stock_news — Stock News
Overview
Financial news linked to individual stocks: headline, summary, source, publish time, and URL.
Files
File
Description
data.parquet
Full news archive
Key Fields
Field
Description
symbol
Associated stock symbol (primary query key)
title
Headline
description
Summary body
publish_date
Publish date (YYYY-MM-DD or locale-specific text)… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_stock_news.telegram-news-ua-dataset
Aisberg Telegram News UA
A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.aria_synthetic_envs_mcmc_3dgs_newLicense Notice:This dataset is derived from the Aria Dataset.It follows the Aria Synthetic Environments Dataset License Agreement.See Aria License for details.
sn80-data-new1new-datasetforecast-news-embeddings
Forecast News Embeddings
Precomputed LanceDB table used for keyword, semantic, and hybrid retrieval in
forecast-sim and future-sim.
Snapshot
7,911,857 indexed source articles
16,207,764 text chunks
Coverage: 2023-01-11 through 2026-08-31
Snapshot published: 2026-09-18
Lance dataset version: 856
Total artifact size: approximately 303.2 GiB
Articles with empty searchable text are not represented. Long articles can
produce multiple chunks, so the chunk count is… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news-embeddings.multi_newsMulti-News, consists of news articles and human-written summaries
of these articles from the site newser.com.
Each summary is professionally written by editors and
includes links to the original articles cited.
There are two features:
- document: text of news articles seperated by special token "|||||".
- summary: news summary.institutional-newspapers-bpl
📰 Institutional Newspapers: Boston Public Library
A structured dataset derived from the Boston Public Library's public domain
newspapers collection, produced by the Institutional Data
Initiative in collaboration with Boston Public Library.
1,473,635 public domain newspaper scans, published between 1795 and 1930
83,147,041 individual crops segmented from those scans
16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract
Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs:
The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date.
We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.infini-news-index
INFINI-NEWS FM-Index
🔎 Live search API: these FM-indexes power a public search service — full-text search, n-gram counts, and document retrieval in the browser or via a keyless REST API, without building the index yourself — at infini-news.uni-graz.at (API reference).
Pre-built FM-indexes (Burrows–Wheeler Transform + suffix array, built
with infini-gram-mini,
Liu et al. 2025) over the
ruggsea/infini-news-corpus
parquets. Enables exact, byte-level substring count and document… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-index.danbooru-tag-csv
danbooru-tag-csv
CSV files of Danbooru tags.
Dataset Description
This project manages CSVs of Danbooru tags, which can be used by Danbooru related applications and libraries.
Dataset Creation
These CSV files were created using the following datasets:
itterative/danbooru_wikis_full
trojblue/danbooru2025-metadata
License
This dataset is released under the MIT License.
US-PD-Newspapers
🇺🇸 US Public Domain Newspapers 🇺🇸
US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library.
With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining.
Content
As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.19c_newspapers_images_altoLeetCodeDataset
LeetCodeDataset
LeetCodeDataset is a dataset consists of Python leetcode problems that can be used for LLM training and evaluation.
💻 GitHub
📄 LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
📄 Policy Filtration for RLHF to Mitigate Noise in Reward Models
stream-data-newMUSE-News
MUSE-News
MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles… See the full description on the dataset page: https://huggingface.co/datasets/muse-bench/MUSE-News.nepali-news-dataset
🇳🇵 Nepali News Dataset & NLP Corpus
The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours.
Repository: thegauravgiri/nepali-news-dataset
Total Articles: 15,000+ full-text articles and growing
Update Frequency: Every 4 hours via automated GitHub Actions pipelines
Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding
License: MIT License
⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.Vietnamese-News
Dataset Card for "VietnameseNewsparquet"
More Information needed
interleaved-umm-new
Interleaved Multimodal Reasoning Dataset
A dataset generation framework for spatial reasoning tasks involving camera viewpoint prediction and ordering around static 3D objects. This project generates multimodal chain-of-thought reasoning traces that teach models how camera views change during orbital rotation.
Overview
This framework generates two types of spatial reasoning tasks:
Task 1: Camera View Prediction - Given an initial view and rotation parameters (angle +… See the full description on the dataset page: https://huggingface.co/datasets/Caesarrr/interleaved-umm-new.ag_newsCodemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
Emotion_new_collected_datasetsenator-tweets
