CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01newfacade /LeetCodeDataset LeetCodeDataset LeetCodeDataset is a dataset consists of Python leetcode problems that can be used for LLM training and evaluation. 💻 GitHub 📄 LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs 📄 Policy Filtration for RLHF to Mitigate Noise in Reward Models texttext-generation1K<n<10K84 likes8.2k downloads1y agoHugging Face02dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face03BAAI /IndustryCorpus_news[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_news.texttext-generation100M<n<1B5 likes1.8k downloads1mo agoHugging Face04common-pile /news News Description We scrape the news sites that publish content under CC BY or CC BY-SA according to opennewswire. These include 360info, Africa is a Country, Alt News, Balkan Diskurs, Factly, Freedom of the Press Foundation, Agenzia Fides, Global Voices, Meduza, Mekong Eye, Milwaukee Neighborhood News Service, Minority Africa, New Canadian Media, SciDev.Net, The Solutions Journalism Exchange, Tasnim News Agency… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/news.texttext-generation100K<n<1M2 likes317 downloads1y agoHugging Face05Alwin-Yang /bmw-pressclub-news BMW PressClub News Dataset This dataset contains press releases and news articles scraped from BMW PressClub. Dataset Structure JSON Format (bmw_articles.json) { "scraped_at": "2025-12-17T10:00:00", "source": "https://www.press.bmwgroup.com/global/article", "count": 100, "articles": [ { "title": "BMW presents the new X5", "date": "17.12.2025", "article_type": "Press Release", "summary": "...", "tags": ["BMW X5", "SUV"]… See the full description on the dataset page: https://huggingface.co/datasets/Alwin-Yang/bmw-pressclub-news.texttext-generationn<1K0 likes275 downloads6mo agoHugging Face06playcat /playcat-cat-behavior-new-data-set PlayCat Cat Behavioral Enrichment Dataset The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research Dataset Summary The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.tabulartext-classification10K<n<100K0 likes131 downloads4mo agoHugging Face07glnmario /news-qa-summarization NewsQASum, a dataset for question answering and summarization of news This dataset contains the CNN articles at the overlap between the newsqa question-answering dataset and the CNN DailyMail summarization dataset. Each article is annotated with a summary and a list of questions and corresponding answers. Tasks: QA, summarization, text retrievalGenre: News storiesLanguage: English textsummarization10K<n<100K31 likes91 downloads3y agoHugging Face08acul3 /KoPI-CC_News Dataset Summary KoPI(Korpus Perayapan Indonesia)-CC_News is Indonesian Only Extract from CC NEWS Common Crawl from 2016-2022(july) ,each snapshots get extracted using warcio,trafilatura and filter using fasttext detail soon texttext-generation1M<n<10M2 likes85 downloads4y agoHugging Face09Jotschi /german-news-titles Dataset Card for German News Titles The dataset contains synthetically generated german news articles and a set of corresponding titles. Dataset Creation The dataset was created using dolphin-mixtral:v2.7. The source scripts generated a news article based on a given topic. For the resulting article multiple titles were generated which are included in the dataset. texttext-generation1K<n<10K2 likes85 downloads2y agoHugging Face10minhhien0811 /ja-current-news-keyword-sft-30k Japanese Current-News Business Keyword SFT 30K Synthetic Japanese SFT dataset for structured keyword generation. Given one theme, the assistant returns JSON with: categories: 3-6 upper-level categories terms: 12-16 related terms each term has label and cat every cat exactly matches one item from categories Themes are current-news oriented and cover topics such as LLMs, AI policy, AI agents, economics, markets, Trump-related policy, tariffs, monetary policy, geopolitics, and… See the full description on the dataset page: https://huggingface.co/datasets/minhhien0811/ja-current-news-keyword-sft-30k.tabulartext-generationn<1K0 likes74 downloads3mo agoHugging Face11jslin09 /news_commentary_tw本資料集是來自QingySi所搜集的中英對照新聞評論,一共有 252,776 對中英語翻譯的句子,是使用Alpaca的指令資料集格式製成。本資料集利用了OpenCC 進行簡轉繁。 texttranslation100K<n<1M3 likes55 downloads3y agoHugging Face12iamsubingyawali /nepali_news_texttexttext-generation100K<n<1M0 likes50 downloads1y agoHugging Face13kurumikz /kaz-news-corpus kaz-news-corpus Қазақ тіліндегі жаңалықтар корпусы — egemen.kz, baq.kz және azattyq.org сайттарынан жиналған. Датасет туралы Параметр Мән Мақала саны 11,814 Жалпы көлемі 53.86 MB Жалпы сөз саны 3,405,903 Орташа мақала ұзындығы 2,274 символ / 288 сөз Медиана ұзындығы 1,503 символ Ең ұзын мақала 81,103 символ Ең қысқа мақала 86 символ Бос тақырып (title) 3 (0.025%) Деректер көздері Сайт Мақала Орташа ұзындық Үлесі Baq.kz 4… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/kaz-news-corpus.texttext-classification10K<n<100K0 likes49 downloads6mo agoHugging Face14Mxode /University-News-Instruction-Zh一些高校校园新闻,约 65k * 3(类任务) 条,稍微做了一点点脱敏,尽可能地遮盖了作者名等。数据已经整理成了指令的形式,格式如下: { "id": <id>, "category": "(title_summarize|news_classify|news_generate)", "instruction": <对应的具体指令>, "input": <空>, "output": <指令对应的输出> } 总共三类任务:标题总结、栏目分类、新闻生成,本质上是利用新闻元数据中的标题、栏目、内容排列组合生成的,所以可以保证数据完全准确。每个字段内容已经整理成了单行的格式。下面是三类任务的样例: // 标题总结 { "id": 22106, "category": "title_summarize", "instruction": "请你给下面的新闻取一则标题:\n点击图片观看视频… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/University-News-Instruction-Zh.textzero-shot-classification100K<n<1M4 likes48 downloads1y agoHugging Face15hausmer /truha-news-satire Truha news satire (Ukrainian) A synthetic, LLM-generated dataset of Ukrainian "news" written in the truha satirical register — short, meta-ironic, punchline-driven fake-news items. The corpus was produced by distilling a target style (a Ukrainian satirical news persona) into generated examples; no real user data is included. Intended use: style-transfer / imitation training for a Ukrainian satirical news-writing assistant. Each example is a (system, instruction, output) triple.… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truha-news-satire.texttext-generation10K<n<100K0 likes45 downloads9d agoHugging Face16pansophic /newsophy-v0.1 This dataset was used to train the pansophic-1-preview model This dataset was created using open-source, permissively licensed models. In addition to providing answers to a diverse set of questions, we leveraged multiple open-source pipelines to generate new tasks and questions, enriching the dataset's variety and complexity. The dataset includes examples that showcase tool usage, contextual understanding, and the application of system prompts. Topics distribtuion in… See the full description on the dataset page: https://huggingface.co/datasets/pansophic/newsophy-v0.1.texttext-generation100K<n<1M1 likes42 downloads1y agoHugging Face17nixjoe /new-cpu-260310 Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw — Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/nixjoe/new-cpu-260310.texttext-generationn<1K0 likes39 downloads7mo agoHugging Face18kurumikz /Kazakh_news_corpus Kazakh Clean News Dataset 🇰🇿 Dataset Description This dataset contains thoroughly cleaned Kazakh language articles collected from leading news and information portals in Kazakhstan. The data was specifically prepared for Natural Language Processing (NLP) tasks, pre-training, and fine-tuning Large Language Models (LLMs) in the Kazakh language. Key Features: All HTML code, advertisements, social media buttons, short news snippets, and subscription prompts have been… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Kazakh_news_corpus.texttext-generationn<1K0 likes30 downloads7mo agoHugging Face19Good-News-Lending /loan-comparison-scenarios Loan Comparison Scenarios 200+ loan comparison scenarios: FHA vs Conv, VA vs FHA, USDA vs FHA, ARM vs Fixed. Details Records: 204 Format: JSONL License: CC-BY-4.0 Last Updated: March 2026 Verified By: Tate Thompson, NMLS #2473962 Publisher: Good News Lending Thompson Alpha Logic Side-by-side program comparisons with deterministic 'Winner' logic based on credit score, down payment, and time horizon. Each scenario shows exact monthly payment, total cost, and… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/loan-comparison-scenarios.tabularquestion-answeringn<1K0 likes30 downloads6mo agoHugging Face20QuantumChem /QuantumChem-200k-new QuantumChem-200K: A Large Molecular Corpus for Chemistry Screening and Discovery. Paper under review. code at: https://github.com/AnonymousUser-3/QuantumChem-200K. license: GLP-3.0 texttext-generation100K<n<1M0 likes30 downloads29d agoHugging Face21newmindai /siu-rag-data SIU-RAG Dataset Overview The newmindai/siu-rag-data dataset is a specialized evaluation dataset designed for benchmarking Retrieval-Augmented Generation (RAG) systems, with a particular focus on analyzing RAG performance with guided decoding methods. This dataset was specifically created for the experiments described in the IEEE paper "Guided Decoding for Retrieval Augmented Generation". Dataset Structure The dataset consists of 507 rows in the training… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/siu-rag-data.texttext-generationn<1K2 likes27 downloads1y agoHugging Face22nixjoe /new-cpu-260308 Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw — Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/nixjoe/new-cpu-260308.texttext-generationn<1K0 likes27 downloads7mo agoHugging Face23TaiMingLu /news-truthfulThis the dataset for Every Language Counts: Learn and Unlearn in Multilingual LLMs. Each of the 100 row contains a GPT generated 'real' news article, a corresponding 'fake' news article with injected fake information, and the 'fake' keyword. It contains 10 Q&A pairs on 'real' news for instruction tunning. We also provide one question to evaluate 'real' news understanding and another question to count the appearance of 'fake' detail. Note: The dataset contains news articles with fake… See the full description on the dataset page: https://huggingface.co/datasets/TaiMingLu/news-truthful.texttext-generationn<1K3 likes26 downloads2y agoHugging Face24itsSHAS /clean_ukrainian-news Ukrainian News Dataset This is a dataset of news articles downloaded from various Ukrainian websites and Telegram channels. The dataset contains 22 567 099 JSON objects (news), total size ~67GB each with the following fields: title: The title of the news article text: The text of the news article, which may contain HTML tags(e.g., paragraphs, links, images, etc.) url: The URL of the news article datetime: The time of publication or when the article was parsed and added to… See the full description on the dataset page: https://huggingface.co/datasets/itsSHAS/clean_ukrainian-news.texttext-generation10M<n<100M0 likes25 downloads8mo agoHugging Face25Good-News-Lending /rent-vs-buy-scenarios Rent Vs Buy Scenarios 2,148 pre-computed rent vs buy breakeven scenarios across 179 markets. Details Records: 2148 Format: JSONL License: CC-BY-4.0 Last Updated: March 2026 Verified By: Tate Thompson, NMLS #2473962 Publisher: Good News Lending Thompson Alpha Logic Economic modeling of the 'Break-Even Year' in 179 Southeast markets based on current rental inflation vs. fixed-rate mortgage stability. Calculates the exact month when buying becomes cheaper than… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/rent-vs-buy-scenarios.tabularquestion-answering1K<n<10K0 likes25 downloads6mo agoHugging Face26PJMixers /AP-News-2024Number of Samples: 1,000 Most Tokens: 4,817 Least Tokens: 186 Average Tokens: 1,027 Median Tokens: 925 Total Tokens: 1,027,462 Using Mistral-v0.1-7B Tokenizer A small set of news articles which I will try to continue collecting. Very recent samples to bring a model up to speed with what's currently happening. Grabbed whatever articles I could find going through different categories. There will be a bias towards whatever AP News has on its front page whenever I open… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers/AP-News-2024.texttext-generation1K<n<10K3 likes24 downloads3y agoHugging Face27Skiittoo /aljazeera-news-arabic Al Jazeera Arabic News Articles Dataset A dataset of 9,758 Arabic news articles scraped from Al Jazeera Arabic (aljazeera.net), covering the period from September 30, 2025 to March 14, 2026. Dataset Description Each record contains the full article text, title, publication date, topic labels, and optional image metadata. The articles span 183 unique topic tags across politics, sports, economy, religion, and more. Supported Tasks Text Classification… See the full description on the dataset page: https://huggingface.co/datasets/Skiittoo/aljazeera-news-arabic.imagetext-classification1K<n<10K0 likes24 downloads6mo agoHugging Face28NewEden-Forge /Orion-Roleplay-Logs-Sharegpt-Ngram-cleanedsame as the previous but filtered "what do you" which was wayyyy too present texttext-generation1K<n<10K3 likes22 downloads2y agoHugging Face29proxectonos /nos-rag-news NOS RAG Dataset (Galician News) Dataset Description This dataset is built around a collection of Galician news articles and a set of question–answer pairs derived from them. The main goal is to provide a compact benchmark for evaluating Retrieval-Augmented Generation (RAG) systems in a realistic setting. Each question is grounded in a specific news article, and answering it correctly requires identifying and using the relevant parts of the source text. In many… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/nos-rag-news.texttext-generationn<1K0 likes22 downloads5mo agoHugging Face30REXX-NEW /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/REXX-NEW/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K6 likes21 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.