CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gfdg34fsd /newe9 likes688k downloads3d agoHugging Face02fancyzhx /ag_news Dataset Card for "ag_news" Dataset Summary AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc)… See the full description on the dataset page: https://huggingface.co/datasets/fancyzhx/ag_news.texttext-classification100K<n<1M195 likes86k downloads3y agoHugging Face03ruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B39 likes45k downloads4d agoHugging Face04open-index /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.texttext-generation10M<n<100M343 likes45k downloads29d agoHugging Face05shash42 /forecast-news Forecast News Deduplicated daily news corpus used by forecast-sim and future-sim. Snapshot 31,859,020 articles 3,463 daily partitions Coverage: 2016-08-26 through 2026-08-31 Snapshot published: 2026-09-18 Stored data size: approximately 158.5 GiB The Parquet files are the canonical complete representation. The repository also contains daily JSONL files where available and compact headline JSON files used by article-browsing workflows. Layout Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.text10M<n<100M2 likes44k downloads3d agoHugging Face06jmhessel /newyorker_caption_contest Dataset Card for New Yorker Caption Contest Benchmarks Dataset Summary See capcon.dev for more! Data from: Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest @inproceedings{hessel2023androids, title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding'' Benchmarks from {The New Yorker Caption Contest}}, author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.imageimage-to-text100K<n<1M76 likes24k downloads3y agoHugging Face07cvlab /new-york-smells New York Smells: A Large Multimodal Dataset for Olfaction While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of diverse, multimodal olfactory data collected in natural settings. We present New York Smells, a large-scale dataset of paired image and olfactory signals captured in-the-wild. Our dataset contains 7,000 smell-image pairs from 3,500 distinct objects… See the full description on the dataset page: https://huggingface.co/datasets/cvlab/new-york-smells.image10K<n<100K1 likes17k downloads2mo agoHugging Face08AlphaDojo /dojo_stock_news Languages: 简体中文 · English dojo_stock_news — Stock News Overview Financial news linked to individual stocks: headline, summary, source, publish time, and URL. Files File Description data.parquet Full news archive Key Fields Field Description symbol Associated stock symbol (primary query key) title Headline description Summary body publish_date Publish date (YYYY-MM-DD or locale-specific text)… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_stock_news.text1M<n<10M0 likes15k downloads47m agoHugging Face09aisbergpublicorganization /telegram-news-ua-dataset Aisberg Telegram News UA A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.text-classification100K<n<1M4 likes14k downloads5m agoHugging Face10ShapeSplats /aria_synthetic_envs_mcmc_3dgs_newgatedLicense Notice:This dataset is derived from the Aria Dataset.It follows the Aria Synthetic Environments Dataset License Agreement.See Aria License for details. 10K<n<100K1 likes14k downloads22d agoHugging Face11dsaddsaf /sn80-data-new10 likes12k downloads1mo agoHugging Face12sagels /new-dataset0 likes11k downloads1m agoHugging Face13shash42 /forecast-news-embeddings Forecast News Embeddings Precomputed LanceDB table used for keyword, semantic, and hybrid retrieval in forecast-sim and future-sim. Snapshot 7,911,857 indexed source articles 16,207,764 text chunks Coverage: 2023-01-11 through 2026-08-31 Snapshot published: 2026-09-18 Lance dataset version: 856 Total artifact size: approximately 303.2 GiB Articles with empty searchable text are not represented. Long articles can produce multiple chunks, so the chunk count is… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news-embeddings.2 likes9.7k downloads3d agoHugging Face14alexfabbri /multi_newsMulti-News, consists of news articles and human-written summaries of these articles from the site newser.com. Each summary is professionally written by editors and includes links to the original articles cited. There are two features: - document: text of news articles seperated by special token "|||||". - summary: news summary.summarization10K<n<100K78 likes9.3k downloads3y agoHugging Face15institutional /institutional-newspapers-bplgated 📰 Institutional Newspapers: Boston Public Library A structured dataset derived from the Boston Public Library's public domain newspapers collection, produced by the Institutional Data Initiative in collaboration with Boston Public Library. 1,473,635 public domain newspaper scans, published between 1795 and 1930 83,147,041 individual crops segmented from those scans 16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.image1M<n<10M9 likes9.2k downloads1mo agoHugging Face16SetFit /20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs: The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date. We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.text10K<n<100K21 likes9k downloads5y agoHugging Face17ruggsea /infini-news-index INFINI-NEWS FM-Index 🔎 Live search API: these FM-indexes power a public search service — full-text search, n-gram counts, and document retrieval in the browser or via a keyless REST API, without building the index yourself — at infini-news.uni-graz.at (API reference). Pre-built FM-indexes (Burrows–Wheeler Transform + suffix array, built with infini-gram-mini, Liu et al. 2025) over the ruggsea/infini-news-corpus parquets. Enables exact, byte-level substring count and document… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-index.text-retrieval2 likes8.7k downloads4d agoHugging Face18newtextdoc1111 /danbooru-tag-csv danbooru-tag-csv CSV files of Danbooru tags. Dataset Description This project manages CSVs of Danbooru tags, which can be used by Danbooru related applications and libraries. Dataset Creation These CSV files were created using the following datasets: itterative/danbooru_wikis_full trojblue/danbooru2025-metadata License This dataset is released under the MIT License. 28 likes8.2k downloads1y agoHugging Face19PleIAs /US-PD-Newspapers 🇺🇸 US Public Domain Newspapers 🇺🇸 US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library. With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining. Content As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.texttext-generation10M<n<100M50 likes8k downloads3y agoHugging Face20ambrosfitz /19c_newspapers_images_altotabular100K<n<1M4 likes7.7k downloads3mo agoHugging Face21newfacade /LeetCodeDataset LeetCodeDataset LeetCodeDataset is a dataset consists of Python leetcode problems that can be used for LLM training and evaluation. 💻 GitHub 📄 LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs 📄 Policy Filtration for RLHF to Mitigate Noise in Reward Models texttext-generation1K<n<10K84 likes7.5k downloads1y agoHugging Face22asdrty123 /stream-data-newaudion<1K1 likes6.8k downloads51m agoHugging Face23muse-bench /MUSE-News MUSE-News MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles… See the full description on the dataset page: https://huggingface.co/datasets/muse-bench/MUSE-News.text10K<n<100K4 likes6.2k downloads2y agoHugging Face24thegauravgiri /nepali-news-dataset 🇳🇵 Nepali News Dataset & NLP Corpus The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours. Repository: thegauravgiri/nepali-news-dataset Total Articles: 15,000+ full-text articles and growing Update Frequency: Every 4 hours via automated GitHub Actions pipelines Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding License: MIT License ⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.texttext-classification10K<n<100K1 likes6.1k downloads3h agoHugging Face25tmnam20 /Vietnamese-News Dataset Card for "VietnameseNewsparquet" More Information needed text1M<n<10M0 likes6k downloads3y agoHugging Face26Caesarrr /interleaved-umm-new Interleaved Multimodal Reasoning Dataset A dataset generation framework for spatial reasoning tasks involving camera viewpoint prediction and ordering around static 3D objects. This project generates multimodal chain-of-thought reasoning traces that teach models how camera views change during orbital rotation. Overview This framework generates two types of spatial reasoning tasks: Task 1: Camera View Prediction - Given an initial view and rotation parameters (angle +… See the full description on the dataset page: https://huggingface.co/datasets/Caesarrr/interleaved-umm-new.0 likes5.5k downloads7mo agoHugging Face27SetFit /ag_newstext100K<n<1M10 likes5.2k downloads5y agoHugging Face28RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes5.2k downloads5mo agoHugging Face29Ayesha758 /Emotion_new_collected_datasetaudio10K<n<100K0 likes5.1k downloads4mo agoHugging Face30m-newhauser /senator-tweetstext10K<n<100K6 likes4.5k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.