CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shash42 /forecast-news Forecast News Deduplicated daily news corpus used by forecast-sim and future-sim. Snapshot 31,859,020 articles 3,463 daily partitions Coverage: 2016-08-26 through 2026-08-31 Snapshot published: 2026-09-18 Stored data size: approximately 158.5 GiB The Parquet files are the canonical complete representation. The repository also contains daily JSONL files where available and compact headline JSON files used by article-browsing workflows. Layout Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.text10M<n<100M2 likes40k downloads4d agoHugging Face02aisbergpublicorganization /telegram-news-ua-dataset Aisberg Telegram News UA A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.texttext-classification100K<n<1M4 likes16k downloads26m agoHugging Face03SetFit /20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs: The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date. We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.text10K<n<100K21 likes9.4k downloads5y agoHugging Face04newfacade /LeetCodeDataset LeetCodeDataset LeetCodeDataset is a dataset consists of Python leetcode problems that can be used for LLM training and evaluation. 💻 GitHub 📄 LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs 📄 Policy Filtration for RLHF to Mitigate Noise in Reward Models texttext-generation1K<n<10K84 likes7.6k downloads1y agoHugging Face05SetFit /ag_newstext100K<n<1M10 likes5.1k downloads5y agoHugging Face06dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4k downloads1y agoHugging Face07Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.40546875 Action score: 0.475 Valid samples: 320/320 tabularn<1K0 likes3.7k downloads14d agoHugging Face08Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4703125 Valid samples: 320/320 tabularn<1K0 likes3.7k downloads14d agoHugging Face09Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.39921875 Action score: 0.44375 Valid samples: 320/320 tabularn<1K0 likes3.7k downloads14d agoHugging Face10Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38359375 Action score: 0.4703125 Valid samples: 320/320 tabularn<1K0 likes3.7k downloads14d agoHugging Face11SetFit /bbc-news BBC News Topic Dataset Dataset on BBC News Topic Classification consisting of 2,225 articles published on the BBC News website corresponding during 2004-2005. Each article is labeled under one of 5 categories: business, entertainment, politics, sport or tech. Original source for this dataset: Derek Greene, Pádraig Cunningham, “Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering,” in Proc. 23rd International Conference on Machine learning (ICML’06)… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/bbc-news.texttext-classification1K<n<10K24 likes2.7k downloads2y agoHugging Face12NewEden /xlam-function-calling-60k-shareGPTShareGPT converted version of Salesforce/xlam-function-calling-60k text10K<n<100K0 likes2.5k downloads2y agoHugging Face13Stage-jh-monitor /qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4 qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3890625 Action score: 0.4359375 Valid samples: 320/320 tabularn<1K0 likes2k downloads8d agoHugging Face14Stage-jh-monitor /qwen35-4b-filter-solvability-200-qwen38-27b-newprompt-4k-epoch4 qwen35-4b-filter-solvability-200-qwen38-27b-newprompt-4k-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.40234375 Action score: 0.421875 Valid samples: 320/320 tabularn<1K0 likes2k downloads8d agoHugging Face15BAAI /IndustryCorpus_news[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_news.texttext-generation100M<n<1B5 likes1.8k downloads1mo agoHugging Face16sh0416 /ag_newsAG's News Topic Classification Dataset Version 3, Updated 09/09/2015 ORIGIN AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search… See the full description on the dataset page: https://huggingface.co/datasets/sh0416/ag_news.texttext-classification100K<n<1M15 likes1.8k downloads4y agoHugging Face17heegyu /news-category-datasetDataset from https://www.kaggle.com/datasets/rmisra/news-category-dataset text100K<n<1M4 likes1.6k downloads4y agoHugging Face18cm2435-new /gdpval_preference_rubricsaudion<1K0 likes1.6k downloads5mo agoHugging Face19titasmallick96 /daily-bio-newsaudion<1K0 likes1.5k downloads4h agoHugging Face20ZechengLi19 /CSL-News Summary This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News is a large-scale Chinese Sign Language dataset designed for developing robust sign language understanding models. Code: https://github.com/ZechengLi19/Uni-Sign Download Please refer to download script to download CSL_News. You can also download each file by wget, for instance: wget… See the full description on the dataset page: https://huggingface.co/datasets/ZechengLi19/CSL-News.textvideo-text-to-text100K<n<1M16 likes1.5k downloads1y agoHugging Face21Sachin21112004 /news-entertainment-datasettexttable-question-answeringn<1K6 likes1k downloads3h agoHugging Face22Askhat777 /GVP_Bot_State_Newtabularn<1K0 likes946 downloads3h agoHugging Face23zuhri025 /IndicVoice-latent-NEWtext100K<n<1M0 likes894 downloads5mo agoHugging Face24robotizac /News_Category_Dataset_v2text100K<n<1M1 likes789 downloads2y agoHugging Face25Sachin21112004 /news-education-datasettext10K<n<100K3 likes773 downloads3h agoHugging Face26Sachin21112004 /news-politics-datasettext1K<n<10K0 likes727 downloads5h agoHugging Face27Sachin21112004 /news-tech-datasettext10K<n<100K1 likes708 downloads3h agoHugging Face28Sachin21112004 /news-finance-datasettext10K<n<100K3 likes707 downloads3h agoHugging Face29jhu-clsp /news21-instructionstexttext-retrieval10K<n<100K1 likes537 downloads6mo agoHugging Face30samsepiol4 /netryx-new-york-5km New York 5km Pre-computed MegaLoc index for Netryx Drishti geolocation. Coverage Center: 40.712800, -74.006000 Radius: 5.0 km Panoramas: 196,824 Index entries: 787,296 Descriptor model: MegaLoc Descriptor dim: 1024 (PCA from 8448) Usage from netryx_hub import NetryxHub hub = NetryxHub() hub.download("new-york-5km", output_dir="./netryx_data/index") # Now open Netryx and search! Or download manually and use Import Index in the Netryx GUI. Details… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-5km.tabularn<1K0 likes492 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.