CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mindweave /web-server-logs Web Server Access Logs (Synthetic) (Free Sample) This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables. Realistic HTTP access logs from a simulated SaaS company running an e-commerce API and marketing website. 50,000 requests across 3 servers over 12 months. Includes realistic patterns: weekday/weekend traffic variation, peak hours, seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and database outage) for anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/web-server-logs.tabulartabular-classification1K<n<10K0 likes1.4k downloads6mo agoHugging Face02APProjects /us-layoffs-monthly-time-series-warn-act US layoffs, month by month — 455 months of WARN notices, 1988-11 → 2026-09, rebuilt daily Last rebuilt: 2026-09-25. One row per calendar month: how many US WARN Act layoff notices were filed, how many workers they named, and how many states contributed — as a regular series with every month present (zeros included), ready for pandas, a chart or a forecasting model. A second table gives the same series per state. 455 consecutive months, 1988-11 → 2026-09, no gaps… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-monthly-time-series-warn-act.tabulartime-series-forecasting1K<n<10K0 likes533 downloads4h agoHugging Face03serenityyyyy /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes507 downloads6mo agoHugging Face04Time-HD-Anonymous /High_Dimensional_Time_Series Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Time-HD-Anonymous/High_Dimensional_Time_Series.tabulartime-series-forecasting100K<n<1M3 likes486 downloads1y agoHugging Face05committa /serena-synthetic-it-28h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.audiotext-to-speech10K<n<100K1 likes415 downloads2mo agoHugging Face06ML-Owl /faang-engineered-time-series-features-2013-2025 FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025) Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!) DOCUMENT NAVIGATION GUIDE (ToC) 1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset 3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.imagetabular-classification10K<n<100K2 likes306 downloads7mo agoHugging Face07sergioburdisso /news_media_bias_and_factuality News Media Factual Reporting and Political Bias Dataset introduced in the paper "Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions" published in the CLEF 2024 main conference. Similar to the news media reliability dataset, this dataset consists of a collections of 4K new media domains names with political bias and factual reporting labels. Columns of the dataset: source: domain name bias: the political bias label. Values: "left"… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_bias_and_factuality.text1K<n<10K4 likes198 downloads2y agoHugging Face08AngieYYF /SPADE-customer-service-dialogue SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection Paper | Code SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B. The datasets are intended for training and evaluating machine generated text detectors in dialogue settings. There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.tabulartext-generation10K<n<100K3 likes187 downloads1y agoHugging Face09Growing-Moss-Data /automotive-service-intelligence-sample 🚗 Automotive Service Intelligence Sample Dataset Connected • Longitudinal • Feature-Engineered • Commercially Available This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development. Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.tabulartabular-classification1K<n<10K2 likes125 downloads3mo agoHugging Face10committa /serena-synthetic-it-27h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.audiotext-to-speech10K<n<100K1 likes122 downloads2mo agoHugging Face11mrcksggcfc /store-sales-time-series-forecasting taken from this Kaggle competition: Dataset Description In this competition, you will predict sales for the thousands of product families sold at Favorita stores located in Ecuador. The training data includes dates, store and product information, whether that item was being promoted, as well as the sales numbers. Additional files include supplementary information that may be useful in building your models. File Descriptions and Data Field Information… See the full description on the dataset page: https://huggingface.co/datasets/mrcksggcfc/store-sales-time-series-forecasting.tabular1M<n<10M3 likes106 downloads15d agoHugging Face12sergiogpinto /memefact-templates MemeFact Templates Dataset This dataset contains 663 meme templates enriched with contextual knowledge for fact-checking meme generation. Each template includes comprehensive information about its origin, cultural significance, visual characteristics, and typical caption patterns to support Retrieval Augmented Generation (RAG) systems. Dataset Description Overview The "MemeFact Templates" dataset is the result of extensive data engineering applied to the… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-templates.imagen<1K1 likes85 downloads1y agoHugging Face13serhanayberkkilic /physiotherapy-evidence-qa 🏥 Physiotherapy Evidence QA: A Bilingual Clinical Corpus Physiotherapy Evidence QA is a large-scale, expert-curated bilingual dataset comprising 143,711 aligned question-answer pairs. It focuses on evidence-based physiotherapy, musculoskeletal rehabilitation, outcome measures, and clinical research methodology. This corpus is designed to facilitate the development of Medical Large Language Models (Med-LLMs), Clinical Decision Support Systems (CDSS), and Cross-Lingual Information… See the full description on the dataset page: https://huggingface.co/datasets/serhanayberkkilic/physiotherapy-evidence-qa.textquestion-answering100K<n<1M4 likes83 downloads10mo agoHugging Face14letrinhan /vn-provinces-retail-trade-and-service-revenue Vietnam provinces retail trade and consumer service revenue Total retail sales of goods and consumer service revenue at current prices (billion VND). Coverage 1995-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Hero (continued) Comparison Color key… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-retail-trade-and-service-revenue.tabular1K<n<10K0 likes83 downloads4d agoHugging Face15jonathansuru /customer_service_intent_detectiontexttext-classification1K<n<10K6 likes69 downloads3y agoHugging Face16sergei202 /nexus-function-callingtextn<1K1 likes69 downloads3y agoHugging Face17FunDialogues /customer-service-robot-support This Dialogue Comprised of fictitious examples of dialogues between a customer encountering problems with a robotic arm and a technical support agent. Check out the example below: "id": 1, "description": "Robotic arm calibration issue", "dialogue": "Customer: My robotic arm seems to be misaligned. It's not picking objects accurately. What can I do? Agent: It appears that the arm may need recalibration. Please follow the instructions in the user manual to reset the calibration… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-robot-support.tabularquestion-answeringn<1K2 likes65 downloads3y agoHugging Face18JLouisBiz /StartYourOwnGoldMine-Sampling-Series-Datasettextn<1K2 likes64 downloads2y agoHugging Face19linaaaaaaaaaaaaa /web-server-logs Web Server Access Logs (Synthetic) (Free Sample) This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables. Realistic HTTP access logs from a simulated SaaS company running an e-commerce API and marketing website. 50,000 requests across 3 servers over 12 months. Includes realistic patterns: weekday/weekend traffic variation, peak hours, seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and database outage) for anomaly… See the full description on the dataset page: https://huggingface.co/datasets/linaaaaaaaaaaaaa/web-server-logs.tabulartabular-classification1K<n<10K1 likes58 downloads3mo agoHugging Face20rkj45 /l2-private-server-openings Lineage 2 Private Server Openings A snapshot of the Lineage 2 private servers tracked by L2 Calendar — a multilingual (EN/ES/PT/RU) calendar and tracker of private server openings by chronicle. Each row is one tracked server with its chronicle, rate, opening date and website. Files l2-private-server-openings.csv — one row per tracked server (480 rows at export time). Columns Column Description name Server name chronicle Lineage 2… See the full description on the dataset page: https://huggingface.co/datasets/rkj45/l2-private-server-openings.tabularn<1K1 likes55 downloads6d agoHugging Face21zaai-ai /time_series_datasets Tourism Monthly Time Series Dataset with Economic and Static Covariates This dataset, originally sourced from Athanasopoulos et al. (2011), focuses on the tourism industry with a monthly frequency and has been enhanced with economic covariates (e.g., CPI, Inflation Rate, GDP) from official Australian government sources. We also perform some preprocessing to further increase the usability of the dataset with dynamic start dates for each series and static covariates for in-depth time… See the full description on the dataset page: https://huggingface.co/datasets/zaai-ai/time_series_datasets.tabulartime-series-forecasting10K<n<100K0 likes54 downloads2y agoHugging Face22sergioburdisso /news_media_reliability Reliability Estimation of News Media Sources: "Birds of a Feather Flock Together" Dataset introduced in the paper "Reliability Estimation of News Media Sources: Birds of a Feather Flock Together" published in the NAACL 2024 main conference. Similar to the news media bias and factual reporting dataset, this dataset consists of a collections of 5.33K new media domains names with reliability labels. Additionally, for some domains, there is also a human-provided reliability score… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_reliability.tabular1K<n<10K2 likes53 downloads2y agoHugging Face23joefee /cell-service-data Dataset Card: Synthetic Mobile Network Performance Dataset Description This dataset contains synthetically generated mobile signal measurements designed to mirror real-world data in the UK. The data represents geolocated signal quality metrics from mobile devices, capturing a range of environmental and temporal conditions over several months in 2025. All data has been anonymized, aggregated, and processed to protect user privacy. The synthetic dataset has undergone pre-processing… See the full description on the dataset page: https://huggingface.co/datasets/joefee/cell-service-data.tabular10M<n<100M3 likes53 downloads1y agoHugging Face24freococo /imam_albani_weaknfab_series_dataset Imam al-Albani Weak & Fabricated Hadith Dataset (AR–EN–MY) This dataset contains weak, rejected, or fabricated hadiths classified byImam Muhammad Nasir al-Din al-Albani, presented in Arabic, English, and Myanmar (Burmese). Translated with Gemini Pro 3.0. Dataset Structure Each row represents one hadith with a global unique ID and multilingual fields. CSV Column Order global_id – Unique sequential ID (primary key) hadith_arabic_text – Original Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/freococo/imam_albani_weaknfab_series_dataset.texttranslation1K<n<10K0 likes50 downloads9mo agoHugging Face25illilliks /multimodal-time-series-forecastingimage100K<n<1M1 likes49 downloads1y agoHugging Face26iBlessi /mcp-server-resource-benchmark MCP Server Resource Benchmark: RAM, Startup, Tool Counts Measured resident memory, startup time and tool counts for 11 popular MCP servers, plus a concurrent five-server stack. Results 178-422MB resident per server (median 188MB) 0.5-1.9 seconds warm startup 961.6MB for a concurrent five-server stack, cross-checked by two independent measurement paths (psutil and PowerShell WorkingSet64) that agreed exactly Runtime floors for context: 52MB bare Node, 15MB bare… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/mcp-server-resource-benchmark.tabularn<1K0 likes48 downloads26d agoHugging Face27pryshlyak /seasonal_time_series_for_anomaly_detection seasonal_time_series_for_anomaly_detection This dataset contains seven CSV files with artificially generated, ordered, timestamped, single-valued metrics for three months divided by days of the week with no anomalies. Also, three CSV files are artificially generated, ordered, timestamped and have single-valued metrics with anomalies, and two CSV files have a week representation (one with anomalies). Motivation This dataset was created as a part of a bachelor's thesis. Our… See the full description on the dataset page: https://huggingface.co/datasets/pryshlyak/seasonal_time_series_for_anomaly_detection.text10K<n<100K1 likes47 downloads2y agoHugging Face28deindexing-services /deindexing-automation-benchmarks Deindexing Automation Engine Benchmarks Benchmark dataset of 20 deindexing automation cases with individual scores for deindex strength, removal rate, review issue handling, reputation health, platform coverage, and workflow efficiency across major removal types and industries. Built by Deindexing.Services. Dataset Description This dataset contains benchmark data for the Deindexing Automation Engine — an automation engine for managing search deindexing, content… See the full description on the dataset page: https://huggingface.co/datasets/deindexing-services/deindexing-automation-benchmarks.tabularn<1K0 likes47 downloads16d agoHugging Face29kyisaiah47 /stacktab-services StackTab: the developer services under price watch One row per developer service whose pricing StackTab reads: its category, its homepage and the pricing page the plan figures were read from. Rows in this cut 29 One row is one service Cut 2026-09-04 Refreshed Monthly, on the first of the month Measured by StackTab Method https://toolproof.thecompound.tech/methodology Licence Creative Commons Attribution 4.0 International Publisher Compound Labs Also… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/stacktab-services.textn<1K0 likes46 downloads6d agoHugging Face30FunDialogues /customer-service-grocery-cashier This Dialogue Comprised of fictitious examples of dialogues between a customer at a grocery store and the cashier. Check out the example below: "id": 1, "description": "Price inquiry", "dialogue": "Customer: Excuse me, could you tell me the price of the apples per pound? Cashier: Certainly! The price for the apples is $1.99 per pound." How to Load Dialogues Loading dialogues can be accomplished using the fun dialogues library or Hugging Face datasets library.… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-grocery-cashier.tabularquestion-answeringn<1K4 likes45 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.