CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B39 likes46k downloads7d agoHugging Face02PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes6.3k downloads3y agoHugging Face03dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face04whiskey1983 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/whiskey1983/hacker-news.tabulartext-generation10M<n<100M0 likes1.6k downloads6mo agoHugging Face05Whiteglove44 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Whiteglove44/hacker-news.tabulartext-generation10M<n<100M0 likes1.6k downloads6mo agoHugging Face06Yahoo-Finance-News /FineWeb2024 FineWeb-Edu 2024 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2024. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2024 Rows 162,500,784… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2024.tabulartext-generation100M<n<1B0 likes1k downloads7d agoHugging Face07biglam /europeana_newspapers Dataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiative, this unofficial version extracts text content from ALTO XML and… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana_newspapers.imagetext-generation10M<n<100M62 likes860 downloads1d agoHugging Face08naveenreddie-18 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/naveenreddie-18/hacker-news.tabulartext-generation10M<n<100M0 likes857 downloads6mo agoHugging Face09Yahoo-Finance-News /FineWeb2025 FineWeb-Edu 2025 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2025 Rows 99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.tabulartext-generation10M<n<100M1 likes745 downloads7d agoHugging Face10nikhilambhure00 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/nikhilambhure00/hacker-news.tabulartext-generation10M<n<100M0 likes646 downloads6mo agoHugging Face11Yahoo-Finance-News /FineWeb-2023 FineWeb-Edu 2023 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2023. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2023 Rows 104,280,950… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb-2023.tabulartext-generation100M<n<1B0 likes620 downloads7d agoHugging Face12MurphyA /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/MurphyA/hacker-news.tabulartext-generation10M<n<100M0 likes525 downloads6mo agoHugging Face13maryamfakhari /crypto-news-coindesk-2020-2025 CoinDesk Cryptocurrency News Dataset (2020–2025) This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics. Time Period January 1, 2020 – January 1, 2025 Content Overview Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.imagetext-classification100K<n<1M2 likes447 downloads1mo agoHugging Face14biglam /hmd_newspapers Dataset Card for Heritage Made Digital Newspapers Dataset Summary This dataset contains text extracted at the article level from historic digitised newspapers from the Heritage Made Digital newspaper digitisation program at the British Library. The newspapers in the dataset were published between 1800 and 1896. This dataset contains ~2.5 billion tokens and 3,065,408 articles. The dataset contains text generated from Optical Character Recognition software on digitised… See the full description on the dataset page: https://huggingface.co/datasets/biglam/hmd_newspapers.tabulartext-generation1M<n<10M10 likes389 downloads3y agoHugging Face15kha2612 /ViSL-News ViSL-News Dataset Summary ViSL-News is a sentence-level Vietnamese Sign Language (VSL) dataset constructed from sign-interpreted Vietnamese news broadcasts. The dataset was built from HTV Tin Tức videos published on YouTube during 2024–2025. Each sample consists of a sentence-level sign-language video clip paired with a Vietnamese text sentence. ViSL-News was constructed using ViSL-Tool, a semi-automated framework designed for news videos that contain spoken… See the full description on the dataset page: https://huggingface.co/datasets/kha2612/ViSL-News.tabulartranslation10K<n<100K0 likes138 downloads17d agoHugging Face16MLDataScientist /Uzbek_news_datasetThis is an Uzbek News Dataset with 512,750 articles (120 million words and in the Latin script) scraped from the web in 2023. I combined and uploaded the dataset in this HF repo so that the community can fine-tune LLMs based on the Uzbek language. @proceedings{kuriyozov_elmurod_2023_7677431, title = {{Text classification dataset and analysis for Uzbek language}}, year = 2023, publisher = {Zenodo}, month = feb, doi =… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/Uzbek_news_dataset.tabulartext-generation100K<n<1M4 likes100 downloads2y agoHugging Face17Fakeepsbiz /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Fakeepsbiz/hacker-news.tabulartext-generation10M<n<100M0 likes100 downloads6mo agoHugging Face18Goader /ukrainian-news-2026 Ukrainian News 2026 Ukrainian-language news articles from 20 national outlets, published between 1 January and 28 August 2026. Extracted body text plus metadata. Two configs. deduplicated is the default — near-duplicates removed, which is what you want when mixing this with an already-deduplicated pretraining corpus. raw is the original release, unchanged. deduplicated (default) raw train-mixin Documents 419,204 429,427 386,477 Characters 0.97B 1.01B 0.88B Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.imagetext-generation1M<n<10M0 likes97 downloads26d agoHugging Face19Linkseed /hacker_news_with_comments Dataset Card for [Dataset Name] Dataset Summary Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags. Supported Tasks and Leaderboards Comment Generation; News analysis with comments; Other comment-based NLP tasks. Languages English Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.tabulartext-generation1M<n<10M6 likes94 downloads4y agoHugging Face20IFthisisrealitynbds /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/IFthisisrealitynbds/hacker-news.tabulartext-generation10M<n<100M0 likes87 downloads6mo agoHugging Face21chrismarchetta /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/chrismarchetta/hacker-news.tabulartext-generation10M<n<100M0 likes75 downloads6mo agoHugging Face22minhhien0811 /ja-current-news-keyword-sft-30k Japanese Current-News Business Keyword SFT 30K Synthetic Japanese SFT dataset for structured keyword generation. Given one theme, the assistant returns JSON with: categories: 3-6 upper-level categories terms: 12-16 related terms each term has label and cat every cat exactly matches one item from categories Themes are current-news oriented and cover topics such as LLMs, AI policy, AI agents, economics, markets, Trump-related policy, tariffs, monetary policy, geopolitics, and… See the full description on the dataset page: https://huggingface.co/datasets/minhhien0811/ja-current-news-keyword-sft-30k.tabulartext-generationn<1K0 likes74 downloads3mo agoHugging Face23biglam /divergent-discourses-tibetan-newspapers Divergent Discourses — Early Tibetan Newspapers, 1950–1965 523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965, produced by the Divergent Discourses project (SOAS University of London and Leipzig University, with Trinity College Dublin). This is not a flat text dump. Each row is one text region from a scanned page, retaining its reading-order position, region type, source newspaper, and issue date — so page structure survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.tabulartext-generation100K<n<1M1 likes72 downloads2mo agoHugging Face24ScoutieAutoML /russian_oil_gas_news_telegram_dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.tabulartext-classification10K<n<100K7 likes62 downloads2y agoHugging Face25TatarNLPWorld /tatar-news-analysis-multilabelgated Dataset Card for Tatar News Multilabel Classification Dataset Details Dataset Description The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.tabulartext-classification10K<n<100K0 likes33 downloads27d agoHugging Face26Good-News-Lending /loan-comparison-scenarios Loan Comparison Scenarios 200+ loan comparison scenarios: FHA vs Conv, VA vs FHA, USDA vs FHA, ARM vs Fixed. Details Records: 204 Format: JSONL License: CC-BY-4.0 Last Updated: March 2026 Verified By: Tate Thompson, NMLS #2473962 Publisher: Good News Lending Thompson Alpha Logic Side-by-side program comparisons with deterministic 'Winner' logic based on credit score, down payment, and time horizon. Each scenario shows exact monthly payment, total cost, and… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/loan-comparison-scenarios.tabularquestion-answeringn<1K0 likes30 downloads6mo agoHugging Face27Good-News-Lending /rent-vs-buy-scenarios Rent Vs Buy Scenarios 2,148 pre-computed rent vs buy breakeven scenarios across 179 markets. Details Records: 2148 Format: JSONL License: CC-BY-4.0 Last Updated: March 2026 Verified By: Tate Thompson, NMLS #2473962 Publisher: Good News Lending Thompson Alpha Logic Economic modeling of the 'Break-Even Year' in 179 Southeast markets based on current rental inflation vs. fixed-rate mortgage stability. Calculates the exact month when buying becomes cheaper than… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/rent-vs-buy-scenarios.tabularquestion-answering1K<n<10K0 likes25 downloads6mo agoHugging Face28lianghsun /tw-news-551Mgated Dataset Card for tw-news-551M tw-news-551M 是一個台灣繁體中文之長文本預訓練語料集,包含 648,576 篇文章,總計約 5.63 億 tokens。涵蓋政治、社會、文化、環境、科技等多元主題,適用於語言模型在繁體中文語境之持續預訓練。 Dataset Details Dataset Description 本資料集彙整台灣繁體中文之公開文章,涵蓋政治、社會、文化、環境、科技等多元主題。每筆資料包含文章全文與 token/字數等統計元資料。 Curated by: Liang Hsun Huang Language(s) (NLP): Traditional Chinese License: CC BY-NC-SA 4.0 Dataset Sources Repository: lianghsun/tw-news-551M Uses Direct… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-news-551M.tabulartext-generation100K<n<1M2 likes23 downloads1mo agoHugging Face29janani-rane /Sinhala-News-Wiki-text-corpus Sinhala-News-Wiki-Text-Corpus Containing news articles from various Sinhala news sites along with Sinhala Wikipedia pages. Dataset Overview Language: Sinhala (සිංහල) Content: Sinhala news articles from various sites Data format: Parquet Number of Records: 18,201 rows (as per current size) Dataset Structure Each record consists of the following fields: category: The news category (e.g., "Other-news, Local-news, wiki, International-news"). site: The site's… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/Sinhala-News-Wiki-text-corpus.tabulartext-classification10K<n<100K0 likes22 downloads2y agoHugging Face30TatarNLPWorld /tatar-news-analysis-multiclassgated Dataset Card for Tatar News Multiclass Classification Dataset Details Dataset Description The Tatar News Multiclass Classification Dataset contains 86,963 Tatar language news articles classified into 9 distinct topic categories. Each entry includes the full article content, title, category (numeric label and text label), source URL, publication date, and content length. The dataset is specifically designed for training and evaluating multi-class… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multiclass.tabulartext-classification10K<n<100K0 likes21 downloads27d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.