headshapeai/infant-head-proportion-measurement-agreement
Photo-Based Measurement Agreement for Infant Head Proportion Aggregate agreement statistics between an automatic photo measurement of infant head proportion and human-reviewed reference values, across 3,349 photographs spanning January 2024 to July 2026. Three rows, 28 columns. No photographs, no per-photograph records, no identifiers — this is aggregate data only. Why this exists We could not find an equivalent figure published by anyone offering this kind of… See the full description on the dataset page: https://huggingface.co/datasets/headshapeai/infant-head-proportion-measurement-agreement.
Photo-Based Measurement Agreement for Infant Head Proportion
Aggregate agreement statistics between an automatic photo measurement of infant head proportion and human-reviewed reference values, across 3,349 photographs spanning January 2024 to July 2026.
Three rows, 28 columns. No photographs, no per-photograph records, no identifiers — this is aggregate data only.
Why this exists
We could not find an equivalent figure published by anyone offering this kind of measurement. Accuracy is asserted often in this space and quantified rarely, and a measurement whose error is unknown is difficult to use well. So we published ours, including the part that does not flatter us: 12.0% of photographs are measured more than five points off.
Headline figures
From the primary window (n = 3,349):
The last row is the one worth carrying. Roughly one photograph in eight is measured poorly, most often where the traced outline picked up bedding, hair or clothing along with the head. An automatic outline-quality check flags 27.5% of photographs and catches 69.9% of those large deviations — useful, but not a substitute for a plain background.
What is measured
Two proportions are read from the traced outline of a head in one photograph taken from directly above, and each is compared against a reference value for the same photograph, reviewed by an experienced team.
Dataset structure
One row per analysis window. All error figures are on the ×100 scale of the quantity concerned (cephalic-index points), not percentages of the value.
The three windows
`primary_2024-05-07_onward` (n = 3,349) — the headline. One photograph per unique image, a single reference convention throughout, every reference value human-reviewed.
`excluded_before_2024-05-07` (n = 910) — published rather than discarded. Excluded from the headline because the reference measurement's own method was still being adjusted in that period: the reference median shifts by 3.0 points as an abrupt step around 2024-05-06/07, while the automatic measurement's own median is identical in both windows (81.9) and so is its segmentation quality (ellipse-residual median 0.0247 in both). The evidence points at the reference side rather than the measurement side — but the exclusion is a judgement call, so its numbers are here for anyone who wants to weigh it.
`all_windows` (n = 4,259) — both combined, published so the effect of the exclusion can be seen rather than taken on trust.
Uses
Suitable for: comparing photo-based anthropometric measurement against a human-reviewed reference; as a reference point for error rates in consumer measurement tools; as an example of a published first-party evaluation including failure rates.
Out-of-scope use
- Not a clinical validation. It compares two measurements of a photograph. It says nothing about health, development, or what any value means for an individual child.
- Not a population reference. This is a convenience sample, not balanced by age and not randomly drawn. The distribution describes the photographs measured, not the prevalence of any proportion among children.
- Reference values were reviewed by an experienced team, not by independent blinded assessors.
- Repeat-photograph consistency — the same head photographed twice, on different devices or in different light — is a separate question and is not addressed here.
Citation
This dataset is archived with a permanent DOI. Cite the DOI, not this page — it resolves for good, and it is the same two files, byte for byte.
doi:10.5281/zenodo.22706467<https://doi.org/10.5281/zenodo.22706467>
Method: <https://www.headshapeai.com/methodology/> Full write-up, including two earlier revisions of this analysis and why they were wrong: <https://www.headshapeai.com/validation/> Benchmark page: <https://www.headshapeai.com/measurement-benchmark/>
Licence
CC BY 4.0 — reuse and redistribution permitted with attribution.
