CoolFace
Datasetpublic

ivitskiy/google-ads-benchmark-2026

Note on checksums. This README.md carries the YAML dataset-card header required by the Hugging Face hub, so its SHA-256 differs from the entry in checksums.txt; that entry refers to the canonical README in the GitHub mirror. All data files are byte-identical across mirrors. Load any table with load_dataset("ivitskiy/google-ads-benchmark-2026", "<config_name>"). Ivitskiy Ads Lab: Google Ads Panel & Benchmark Compilation 2026 (Open Research Dataset) Two things in one package.… See the full description on the dataset page: https://huggingface.co/datasets/ivitskiy/google-ads-benchmark-2026.

sourceHugging Facecc-by-4.0updated 2d agoView on Hugging Face
0likes44downloads
Dataset Card
Note on checksums. This README.md carries the YAML dataset-card header required by the Hugging Face hub, so its SHA-256 differs from the entry in checksums.txt; that entry refers to the canonical README in the GitHub mirror. All data files are byte-identical across mirrors. Load any table with load_dataset("ivitskiy/google-ads-benchmark-2026", "<config_name>").

Ivitskiy Ads Lab: Google Ads Panel & Benchmark Compilation 2026 (Open Research Dataset)

Two things in one package. First, an aggregated, k-anonymous panel of real managed Google Ads accounts (124 spend-active advertisers from one MCC, 2018 to mid-2026, plus 33 audited accounts), publishing two distributions absent from the public record: the Smart Bidding target achievement gap (achieved vs. a July 2026 target snapshot) and the tracked wasted spend rate (share of visible search-terms spend on queries with zero attributed conversions). Second, an industry benchmark compilation: 1,516 published values from 12 publisher families in one machine-readable table, every value carrying its source URL and exact quote.

Requested citation:

Ivitskiy, I. (2026). Ivitskiy Ads Lab Google Ads Panel & Benchmark Compilation 2026 (Version 1.0) [Data set]. Hugging Face. https://huggingface.co/datasets/ivitskiy/google-ads-benchmark-2026

License: CC BY 4.0 (attribution required). Author: Igor Ivitskiy, PhD (ORCID 0000-0002-9749-6414).

Why this dataset exists

Industry benchmark reports (WordStream/LocaliQ and others) answer "what is a normal CTR or CPC". Two questions they do not answer:

  1. 1.How far do Smart Bidding results land from their targets? Advertisers set target CPA / target ROAS. Nobody publishes the distribution of how far actual results land from those targets. (This measures the target achievement gap; it cannot by itself separate algorithm underperformance from advertiser target-setting drift.)
  2. 2.How much search budget is wasted? The share of spend going to search terms that produced zero conversions is the core number behind every "audit your account" pitch, yet no one publishes its distribution.

This dataset publishes both, computed from real managed accounts, alongside conventional CTR/CPC benchmarks and a sourced compilation of the public industry numbers.

Files

FileWhat it contains
data/smart_bidding_promise_gap.csvDistribution of achieved/target ratios for tCPA and tROAS campaigns, Jan to May 2026 (campaign-level)
data/smart_bidding_promise_gap_advertiser_level.csvSame question with the advertiser as the unit of analysis (median of each advertiser's campaigns)
data/smart_bidding_promise_gap_by_vertical.csvCampaign-level gap split by vertical where n >= 7 holds
data/wasted_spend_rate.csvShare of tracked search-terms spend with zero conversions, windows 2023 to 2025
data/panel_ctr_cpc_by_vertical_year.csvMedian (and, for large cells, quartile) CTR / CPC of this panel, search campaigns, vertical x year. Panel levels, not global norms: the panel is Eastern-Europe-heavy; for US-style levels use the industry meta table
data/panel_ctr_cpc_seasonality.csvPanel CTR / CPC medians by calendar month (2022 to 2026H1 pool); same panel caveat
data/panel_device_mix_search.csvMedian device spend shares by vertical (2024 to 2026H1); same panel caveat
data/industry_meta_benchmark.csvSource-attributed compilation of published Google Ads benchmark numbers with URL and exact quote for every value; year is the year as stated by the source (report year and data year can differ, see quote). See the meta-benchmark note below for exact coverage
data/DICTIONARY.mdColumn-level data dictionary
etl/dataset_build.pyThe aggregation code that produced every data/*.csv from the private warehouse

Headline findings (v1)

  • Individual Smart Bidding campaigns land far from their targets much more often than advertiser portfolios do. Measured against a July 2026 target snapshot (within-window target changes are not observable, see Methodology), in Jan to May 2026 33.5% of tCPA campaigns in this panel exceeded their target CPA by more than 20% and 47.3% landed at or under target (tROAS campaigns: 31.3% and 56.3%). At the advertiser level the picture is calmer: 60.9% of tCPA advertisers landed at or under target across their portfolio (median achieved-to-target ratio 0.977). Portfolio averaging absorbs large campaign-level swings; both units of analysis are published so readers can quote either, with the snapshot caveat attached.
  • The median advertiser in this panel spent 39% of tracked search-terms cost on queries with zero attributed conversions (Q1 28%, Q3 67%; windows 2023 to 2025; entities with USD 500+ tracked spend). Spend-weighted, the pooled share is 22.3%: larger budgets show proportionally lower zero-conversion share, whatever the cause. Tracked terms cover about 67% of total search spend in this panel; Google hides the rest below privacy thresholds, and the hidden tail likely converts worse, so these are shares of visible spend and a probable lower bound.
  • Panel CTR / CPC medians by vertical and month are included for continuity with existing public benchmarks (panel levels, not global norms), and industry_meta_benchmark.csv compiles 1,516 published values from 12 publisher families: the full WordStream/LocaliQ annual series (12 report editions, 2016 to 2026, the bulk of the rows) plus Databox, Store Growers, First Page Sage, Wolfgang Digital, Optmyzr, SISTRIX, Adthena, Varos (archived pages) and official Google statistics. Every row carries the source URL and exact quote; 30 headline values were independently re-verified against the source pages during compilation.

Spend glossary (three different scopes; they are consistent, not contradictory): USD 16.8M = total managed spend of the 124-account MCC panel, 2018 to 2026. USD 85M = tracked search-terms spend across BOTH samples: the MCC panel plus the 33 separately audited accounts, several of which are large international SaaS advertisers whose spend is not part of the 16.8M MCC figure (the waste metric uses the 2023 to 2025 subset of this). USD 0.85M = spend covered by the Smart Bidding gap slice (five months, campaigns with active targets and 5+ conversions).

Methodology and validity rules

Sample. Two samples of real Google Ads accounts under management or audit by Ivitskiy Ads Lab: (a) 124 spend-active advertiser accounts from a single MCC, monthly data 2018-01 to 2026-07; (b) 33 audited accounts (snapshot windows 2023 to 2026). Combined tracked search-terms spend: USD 85M. Geography skews to Eastern Europe and international SaaS; this is a convenience sample, not a representative panel. For representative averages use the industry meta-benchmark table; the value of the primary sample is in metrics nobody else publishes, where any real distribution beats no data.

Anonymity (layered thresholds). Every published cell aggregates at least 5 distinct advertiser entities; accounts belonging to one client are collapsed into a single entity before counting ("family dedup"). Because quantiles of small samples reproduce individual values, quartiles are published only for cells with at least 15 entities and spend-weighted (pooled) ratios only for cells with at least 10; smaller cells carry medians only. Per-cell advertiser counts are withheld (a single per-vertical count over the whole window is given instead) to prevent temporal differencing attacks, and per-vertical spend totals are not published to protect dominant spenders. No account identifiers, client names, or domains appear anywhere in the data. Verticals are coarse classes. Search terms were PII-masked (emails, phone numbers, long digit strings) before any processing.

Why there are no cross-advertiser CPA / ROAS / CVR benchmarks. Every account tracks different conversion events (purchase vs. lead vs. trial). Averaging CPA across advertisers whose "conversion" means different things produces numbers that look precise and mean nothing. We publish cross-advertiser stats only for metrics that are comparable by construction: CTR, CPC, spend shares, and within-campaign ratios. This is a deliberate departure from common practice.

Smart Bidding target achievement gap. For every campaign with an active target CPA or target ROAS and at least 5 conversions in the window: miss_ratio = achieved CPA / target CPA (for tROAS: target / achieved), so values above 1.0 mean the campaign did worse than its target. The window is deliberately narrow: January to May 2026 only. Targets are a July 2026 snapshot, and advertisers adjust targets over time, so applying current targets to older performance would manufacture fictional gaps; June 2026 is excluded to let conversions mature for at least 25 days before the data pull. Honest limitations, stated plainly: (a) targets may still have moved within the five-month window and the Google Ads API does not expose target history, so the share of moved targets is not measurable in this release (per-quarter target snapshots are planned from v2 so future releases can report constant-target gaps); (b) the metric measures achievement of the advertiser's chosen target and cannot attribute a miss to the algorithm vs. an unrealistic target; (c) the 5+ conversions filter skews the sample toward campaigns that work at all.

Wasted spend rate. Per advertiser entity: share of tracked search-terms cost (USD) on terms with zero attributed conversions, windows 2023 to 2025 (the 2026 window is excluded because recent conversions have not matured). Both samples are pooled; each entity's costs are summed across its available windows before the share is computed. Conversions are as attributed by each account's own Google Ads conversion settings at pull time (a mix of attribution models across accounts, predominantly account defaults); "zero attributed conversions" is therefore not identical to "zero actual conversions". Entities need at least USD 500 of tracked terms spend. Important scope note: Google reports search terms only above privacy thresholds; tracked terms cover roughly 67% of total search spend in our panel, and the hidden long tail converts worse than average, so true waste is likely higher than these figures. The metric is named trackedwasteshare accordingly.

Currency. All monetary values normalized to USD at fixed mid-2026 rates (documented in the build script). Year-over-year CPC comparisons for non-USD accounts therefore contain exchange-rate artifacts; historical monthly FX is planned for v1.1.

Unit of analysis. The target-achievement-gap tables (file names keep the historical smart_bidding_promise_gap prefix) exist at two levels because they answer different questions. Campaign-level rows (239 tCPA campaigns from 23 advertisers) describe how often an individual Smart Bidding campaign lands near its target; multi-campaign advertisers weigh more there. Advertiser-level rows collapse each advertiser to the median of their campaigns first. The 5+ conversions filter also means the sample skews toward campaigns that work at all; completely failing campaigns are underrepresented. Quote the level you mean.

Reproducibility. The full aggregation code that produced every published CSV ships in this repository as etl/dataset_build.py (reads a private DuckDB warehouse of pseudonymized accounts; all thresholds and definitions are in the code). Raw account-level data is not published, by design.

Applicability and limits → docs/

See Applicability for each file's units, coverage, selection rules and gaps; Sensitivity and Concentration for the advertiser LOO and spend-concentration audit; Reproduce for commands, hashes and missing build inputs; and How to cite for EN/UK/RU examples. The released CSVs contain no advertiser IDs or individual contributions, so LOO stability and advertiser spend concentration cannot be calculated from them. The applicability notes also identify differences between the documented periods/network scope and the shipped ETL; consult them before quoting a number. These documents prepare the dataset for the owner's publication decision.

Privacy statement

This dataset publishes only aggregated statistics behind layered anonymity thresholds (5+ entities per cell for medians, 10+ for spend-weighted ratios, 15+ for quartiles, entity-level deduplication, no per-cell counts, no per-vertical spend totals). Individual account identity is not recoverable by outsiders from the published cells. One honest residual: an advertiser who knows they belong to a small vertical peer group in this panel may recognize that a published median describes their peer group; no identities, values-to-identity mappings, or domains are recoverable even then.

Where to get it

  • Canonical page (tables, context, four headline CSVs): https://ivitskiy.com/blog/en/google-ads-benchmarks-2026/
  • Full dataset, machine-readable, load_dataset()-ready: https://huggingface.co/datasets/ivitskiy/google-ads-benchmark-2026
  • Source mirror with ETL, EDA and docs: https://github.com/ivitskiy/google-ads-benchmark-2026

Versioning and updates

Version 1.0 covers panel data through July 2026 (target snapshot July 2026; waste windows 2023 to 2025). Releases are versioned and every change is logged in CHANGELOG.md. There is no fixed update schedule: a new version is published when the panel or the compilation is materially extended.

Author

Igor Ivitskiy, PhD in mathematical modeling; Google Ads practitioner since 2012. More: ivitskiy.com | ORCID 0000-0002-9749-6414 | LinkedIn

Questions, corrections, collaboration: open an issue in this repository.