CoolFace
Datasetpublic

RotgarSett/chatgpt-clinic-shortlists-model-stability

Clinic Shortlists and Sources Across ChatGPT Models and Reasoning Efforts Version 1.0 of a descriptive stability benchmark of clinic shortlists and reader-visible sources across ChatGPT model and reasoning-effort configurations. Author: Evgeniy Yudin, Founder and Strategy Lead, Rotgar Research Version DOI: 10.5281/zenodo.22162725 Research article and methodological context: rotgar.com License: CC BY 4.0 Published: 2026-08-29 Scope The core benchmark contains 450… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/chatgpt-clinic-shortlists-model-stability.

sourceHugging Facecc-by-4.0updated 28d agoView on Hugging Face
0likes106downloads
Dataset Card

Clinic Shortlists and Sources Across ChatGPT Models and Reasoning Efforts

Version 1.0 of a descriptive stability benchmark of clinic shortlists and reader-visible sources across ChatGPT model and reasoning-effort configurations.

Scope

The core benchmark contains 450 answers: three exact NYC commercial prompts across five ChatGPT model/reasoning configurations, with 30 isolated executions per cell. Fifteen technical preflight answers are excluded from all reported denominators.

Headline within-configuration provider-set overlap was 80.3% for IVF, 34.7% for full-arch dental implants, and 53.9% for bariatric programs. Corresponding reader-visible domain overlap was 51.1%, 28.3%, and 44.8%.

The reader-visible-source core contains 3,757 markdown-link occurrences, 3,690 unique answer-by-URL pairs, and 3,194 answer-by-domain pairs.

Methodological boundary

Collection used Codex CLI 0.147.0 authenticated through one ChatGPT login with live search. It did not use the consumer ChatGPT Free/Plus UI or API-key billing. Position means presentation order, not quality ranking. Reader-visible domains do not establish causal recommendation sources. Physical US IP/location was not controlled.

Boris Teplyakov reviewed the methodology as SEO Lead. This was not clinical review, peer review, a formal external audit, or an assessment of named providers.

Files

The byte-identical public package is mirrored in `package-v1.0/`. Package checksums are listed in `package-v1.0/SHA256SUMS.txt`.

Citation

Yudin, E. (2026). Clinic Shortlists and Sources Across ChatGPT Models and Reasoning Efforts (Version 1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.22162725