datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
web-server-logs
Web Server Access Logs (Synthetic) (Free Sample)
This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables.
Realistic HTTP access logs from a simulated SaaS company running an
e-commerce API and marketing website. 50,000 requests across 3 servers
over 12 months.
Includes realistic patterns: weekday/weekend traffic variation, peak hours,
seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and
database outage) for anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/web-server-logs.fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.High_Dimensional_Time_Series
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Time-HD-Anonymous/High_Dimensional_Time_Series.us-layoffs-monthly-time-series-warn-act
US layoffs, month by month — 455 months of WARN notices, 1988-11 → 2026-09, rebuilt daily
Last rebuilt: 2026-09-17. One row per calendar month: how many US WARN Act layoff
notices were filed, how many workers they named, and how many states contributed —
as a regular series with every month present (zeros included), ready for pandas,
a chart or a forecasting model. A second table gives the same series per state.
455
consecutive months, 1988-11 → 2026-09, no gaps… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-monthly-time-series-warn-act.serena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.faang-engineered-time-series-features-2013-2025
FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025)
Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!)
DOCUMENT NAVIGATION GUIDE (ToC)
1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset
3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.news_media_bias_and_factuality
News Media Factual Reporting and Political Bias
Dataset introduced in the paper "Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions" published in the CLEF 2024 main conference.
Similar to the news media reliability dataset, this dataset consists of a collections of 4K new media domains names with political bias and factual reporting labels.
Columns of the dataset:
source: domain name
bias: the political bias label. Values: "left"… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_bias_and_factuality.SPADE-customer-service-dialogue
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
Paper | Code
SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.automotive-service-intelligence-sample
🚗 Automotive Service Intelligence Sample Dataset
Connected • Longitudinal • Feature-Engineered • Commercially Available
This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development.
Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.serena-synthetic-it-27h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.store-sales-time-series-forecasting
taken from this Kaggle competition:
Dataset Description
In this competition, you will predict sales for the thousands of product families sold at Favorita stores located in Ecuador. The training data includes dates, store and product information, whether that item was being promoted, as well as the sales numbers. Additional files include supplementary information that may be useful in building your models.
File Descriptions and Data Field Information… See the full description on the dataset page: https://huggingface.co/datasets/mrcksggcfc/store-sales-time-series-forecasting.memefact-templates
MemeFact Templates Dataset
This dataset contains 663 meme templates enriched with contextual knowledge for fact-checking meme generation. Each template includes comprehensive information about its origin, cultural significance, visual characteristics, and typical caption patterns to support Retrieval Augmented Generation (RAG) systems.
Dataset Description
Overview
The "MemeFact Templates" dataset is the result of extensive data engineering applied to the… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-templates.vn-provinces-retail-trade-and-service-revenue
Vietnam provinces retail trade and consumer service revenue
Total retail sales of goods and consumer service revenue at current prices (billion VND). Coverage 1995-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-retail-trade-and-service-revenue.StartYourOwnGoldMine-Sampling-Series-Datasetcustomer_service_intent_detectionnexus-function-callingcustomer-service-robot-support
This Dialogue
Comprised of fictitious examples of dialogues between a customer encountering problems with a robotic arm and a technical support agent. Check out the example below:
"id": 1,
"description": "Robotic arm calibration issue",
"dialogue": "Customer: My robotic arm seems to be misaligned. It's not picking objects accurately. What can I do? Agent: It appears that the arm may need recalibration. Please follow the instructions in the user manual to reset the calibration… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-robot-support.novaserve-service-sites-81795545
NovaServe Service-Sites Registry
This dataset contains the service-site registry for NovaServe Facilities Management Pte Ltd, Singapore.
Description
Each row represents a commercial building under a NovaServe maintenance contract, including the site
identifier, building name, street address, planning district, the date of the most recent maintenance
visit, the current maintenance status, and the service-contract tier.
Columns
site_id: unique site… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/novaserve-service-sites-81795545.physiotherapy-evidence-qa
🏥 Physiotherapy Evidence QA: A Bilingual Clinical Corpus
Physiotherapy Evidence QA is a large-scale, expert-curated bilingual dataset comprising 143,711 aligned question-answer pairs. It focuses on evidence-based physiotherapy, musculoskeletal rehabilitation, outcome measures, and clinical research methodology.
This corpus is designed to facilitate the development of Medical Large Language Models (Med-LLMs), Clinical Decision Support Systems (CDSS), and Cross-Lingual Information… See the full description on the dataset page: https://huggingface.co/datasets/serhanayberkkilic/physiotherapy-evidence-qa.time_series_datasets
Tourism Monthly Time Series Dataset with Economic and Static Covariates
This dataset, originally sourced from Athanasopoulos et al. (2011), focuses on the tourism industry with a monthly frequency and has been enhanced with economic covariates (e.g., CPI, Inflation Rate, GDP) from official Australian government sources. We also perform some preprocessing to further increase the usability of the dataset with dynamic start dates for each series and static covariates for in-depth time… See the full description on the dataset page: https://huggingface.co/datasets/zaai-ai/time_series_datasets.multimodal-time-series-forecastingweb-server-logs
Web Server Access Logs (Synthetic) (Free Sample)
This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables.
Realistic HTTP access logs from a simulated SaaS company running an
e-commerce API and marketing website. 50,000 requests across 3 servers
over 12 months.
Includes realistic patterns: weekday/weekend traffic variation, peak hours,
seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and
database outage) for anomaly… See the full description on the dataset page: https://huggingface.co/datasets/linaaaaaaaaaaaaa/web-server-logs.cell-service-data
Dataset Card: Synthetic Mobile Network Performance
Dataset Description
This dataset contains synthetically generated mobile signal measurements designed to mirror real-world data in the UK. The data represents geolocated signal quality metrics from mobile devices, capturing a range of environmental and temporal conditions over several months in 2025.
All data has been anonymized, aggregated, and processed to protect user privacy. The synthetic dataset has undergone pre-processing… See the full description on the dataset page: https://huggingface.co/datasets/joefee/cell-service-data.seasonal_time_series_for_anomaly_detection
seasonal_time_series_for_anomaly_detection
This dataset contains seven CSV files with artificially generated, ordered, timestamped, single-valued metrics for three months divided by days of the week with no anomalies. Also, three CSV files are artificially generated, ordered, timestamped and have single-valued metrics with anomalies, and two CSV files have a week representation (one with anomalies).
Motivation
This dataset was created as a part of a bachelor's thesis. Our… See the full description on the dataset page: https://huggingface.co/datasets/pryshlyak/seasonal_time_series_for_anomaly_detection.imam_albani_weaknfab_series_dataset
Imam al-Albani Weak & Fabricated Hadith Dataset (AR–EN–MY)
This dataset contains weak, rejected, or fabricated hadiths classified byImam Muhammad Nasir al-Din al-Albani, presented in Arabic, English, and Myanmar (Burmese). Translated with Gemini Pro 3.0.
Dataset Structure
Each row represents one hadith with a global unique ID and multilingual fields.
CSV Column Order
global_id – Unique sequential ID (primary key)
hadith_arabic_text – Original Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/freococo/imam_albani_weaknfab_series_dataset.mcp-server-resource-benchmark
MCP Server Resource Benchmark: RAM, Startup, Tool Counts
Measured resident memory, startup time and tool counts for 11 popular MCP servers, plus a concurrent five-server stack.
Results
178-422MB resident per server (median 188MB)
0.5-1.9 seconds warm startup
961.6MB for a concurrent five-server stack, cross-checked by two independent measurement paths (psutil and PowerShell WorkingSet64) that agreed exactly
Runtime floors for context: 52MB bare Node, 15MB bare… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/mcp-server-resource-benchmark.deindexing-automation-benchmarks
Deindexing Automation Engine Benchmarks
Benchmark dataset of 20 deindexing automation cases with individual scores for deindex strength, removal rate, review issue handling, reputation health, platform coverage, and workflow efficiency across major removal types and industries.
Built by Deindexing.Services.
Dataset Description
This dataset contains benchmark data for the Deindexing Automation Engine — an automation engine for managing search deindexing, content… See the full description on the dataset page: https://huggingface.co/datasets/deindexing-services/deindexing-automation-benchmarks.news_media_reliability
Reliability Estimation of News Media Sources: "Birds of a Feather Flock Together"
Dataset introduced in the paper "Reliability Estimation of News Media Sources: Birds of a Feather Flock Together" published in the NAACL 2024 main conference.
Similar to the news media bias and factual reporting dataset, this dataset consists of a collections of 5.33K new media domains names with reliability labels. Additionally, for some domains, there is also a human-provided reliability score… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_reliability.customer-service-grocery-cashier
This Dialogue
Comprised of fictitious examples of dialogues between a customer at a grocery store and the cashier. Check out the example below:
"id": 1,
"description": "Price inquiry",
"dialogue": "Customer: Excuse me, could you tell me the price of the apples per pound? Cashier: Certainly! The price for the apples is $1.99 per pound."
How to Load Dialogues
Loading dialogues can be accomplished using the fun dialogues library or Hugging Face datasets library.… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-grocery-cashier.Telangana_time_series_2023-2025The dataset was retrieved from Open Data Telangana, from February 1, 2023, to January 31, 2025 with daily granularity. The dataset contains various fields such as District, Mandal, Date, rainfall (in millimeters), minimum and maximum temperature (in Celsius), minimum and maximum wind speed, and humidity. It provides a District and Mandal wise distribution as well.
Total Rows - 4,45,213
Total Columns - 10
