CoolFace
Datasetpublic

RotgarSett/medical-google-chatgpt-gemini-source-overlap

Google and AI Source Overlap Across 12 Medical Niches An open, reproducible US dataset comparing explicit ChatGPT and Gemini citations with paired Google organic Top 20 results across 12 medical niches and 432 frozen questions. Full study: https://rotgar.com/medical/resources/google-top-20-chatgpt-gemini-source-overlap Version DOI: https://doi.org/10.5281/zenodo.21850734 Version: 1.0 Fieldwork: August 7, 2026 Publication date: August 8, 2026 Market and language: United States… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/medical-google-chatgpt-gemini-source-overlap.

sourceHugging Facecc-by-4.0updated 28d agoView on Hugging Face
0likes28downloads
Dataset Card

Google and AI Source Overlap Across 12 Medical Niches

An open, reproducible US dataset comparing explicit ChatGPT and Gemini citations with paired Google organic Top 20 results across 12 medical niches and 432 frozen questions.

  • —Full study: https://rotgar.com/medical/resources/google-top-20-chatgpt-gemini-source-overlap
  • —Version DOI: https://doi.org/10.5281/zenodo.21850734
  • —Version: 1.0
  • —Fieldwork: August 7, 2026
  • —Publication date: August 8, 2026
  • —Market and language: United States, English
  • —License: CC BY 4.0
  • —Author: Evgeniy Yudin
  • —ORCID: https://orcid.org/0009-0007-8400-9561
  • —Methodology reviewer: Boris Teplyakov, SEO Lead

Boris Teplyakov reviewed the research methodology only. This was not a medical subject-matter review of the questions, model answers, or cited sources.

Headline common-support result

Among the 322 tasks with comparable observations for both AI platforms, the mean share of cited URLs also found in the paired Google organic Top 20 was 13.6% for ChatGPT and 53.9% for Gemini.

These are descriptive overlap estimates. They do not establish that Google inclusion causes an AI citation, and they are not measures of source quality, trust, answer accuracy, rankings, traffic, leads, conversions, or patient outcomes.

Dataset structure

The root card exposes seven CSV subsets in the Hugging Face Dataset Viewer. The complete 19-file release is preserved under package-v1.0/, including CSV, JSON, XLSX, the frozen question manifest, data dictionary, metadata, citation files, SVG/PNG assets, manifest, license, and checksums.

Viewer subsetFilePurpose
summarymedical-source-overlap-v1.0.csvSector and niche summary distributions
task_metricstask-metrics.csv735 task-platform observations used in released Top 20 estimates
question_manifestquestion-manifest.csv432 frozen questions and design fields
ai_domain_jaccardtask-ai-domain-jaccard.csv336 within-task ChatGPT–Gemini domain-similarity observations
top_cited_domainstop-cited-domains.csvDescriptive frequency table of commonly cited domains
top_matched_domainstop-matched-domains.csvCitation domains also observed in paired Google Top 20
data_dictionarydata-dictionary.csvField definitions

Run sha256sum -c SHA256SUMS.txt inside package-v1.0/ to verify the immutable package files.

Measurement

Each frozen question pairs Google organic results with explicit ChatGPT and Gemini citations. The primary exact-URL metric is the within-task share of unique normalized AI-cited URLs also present in the paired Google organic Top 20. The domain metric repeats the calculation at registrable-domain level. Sector and niche estimates are unweighted macro means across included task-platform observations.

Confidence intervals use 10,000 task-clustered percentile bootstrap replicates with seed 20260807. The common-support comparison uses the same included task set for both AI surfaces. ChatGPT–Gemini domain Jaccard is calculated within task before aggregation.

Design and interpretation limitations

  • —The study is descriptive, not causal.
  • —Source overlap is not a measure of accuracy, quality, endorsement, visibility value, traffic, or conversion.
  • —In the shared 20-niche study program, the first four niches were collected earlier through Live API and the next 16 later through Standard queue. Cross-niche comparisons are partly mixed with collection time and mode.
  • —Extension outcomes were frozen before extension analysis, but the extension was designed after the original four-niche results had been viewed. The full 20-niche panel is therefore post-outcome exploratory.
  • —Every extension row in the released v1.0 exports carries include_primary=true. Treat this as an inclusion flag for the released estimates, not as evidence that the entire 20-niche panel was specified before outcome review.
  • —Questions and model answers did not receive licensed medical subject-matter review.
  • —The package contains derived research data and documentation, not full model answers or complete third-party payloads.

Reuse and citation

The data are available under CC BY 4.0. Attribute the author, identify version 1.0, and cite the version DOI:

bibtex
@dataset{yudin_2026_medical_source_overlap,
  author    = {Yudin, Evgeniy},
  title     = {Google and AI Source Overlap Across 12 Medical Niches},
  year      = {2026},
  publisher = {Zenodo},
  version   = {1.0},
  doi       = {10.5281/zenodo.21850734},
  url       = {https://doi.org/10.5281/zenodo.21850734}
}

The package also includes CITATION.cff and citation.bib. For this published version, use the version DOI above.