CoolFace
Datasetpublic

headshapeai/infant-head-proportion-measurement-agreement

Photo-Based Measurement Agreement for Infant Head Proportion Aggregate agreement statistics between an automatic photo measurement of infant head proportion and human-reviewed reference values, across 3,349 photographs spanning January 2024 to July 2026. Three rows, 28 columns. No photographs, no per-photograph records, no identifiers — this is aggregate data only. Why this exists We could not find an equivalent figure published by anyone offering this kind of… See the full description on the dataset page: https://huggingface.co/datasets/headshapeai/infant-head-proportion-measurement-agreement.

sourceHugging Facecc-by-4.0updated 13d agoView on Hugging Face
0likes42downloads
Dataset Card

Photo-Based Measurement Agreement for Infant Head Proportion

Aggregate agreement statistics between an automatic photo measurement of infant head proportion and human-reviewed reference values, across 3,349 photographs spanning January 2024 to July 2026.

Three rows, 28 columns. No photographs, no per-photograph records, no identifiers — this is aggregate data only.

Why this exists

We could not find an equivalent figure published by anyone offering this kind of measurement. Accuracy is asserted often in this space and quantified rarely, and a measurement whose error is unknown is difficult to use well. So we published ours, including the part that does not flatter us: 12.0% of photographs are measured more than five points off.

Headline figures

From the primary window (n = 3,349):

MetricValue
Three-way agreement (long / typical / wide)85.9%
Mean absolute error, width-to-length2.46 points
Median absolute error0.56 points
Recall on clearly visible left–right differences80.5%
Photographs off by more than 5 points12.0%

The last row is the one worth carrying. Roughly one photograph in eight is measured poorly, most often where the traced outline picked up bedding, hair or clothing along with the head. An automatic outline-quality check flags 27.5% of photographs and catches 69.9% of those large deviations — useful, but not a substitute for a plain background.

What is measured

Two proportions are read from the traced outline of a head in one photograph taken from directly above, and each is compared against a reference value for the same photograph, reviewed by an experienced team.

QuantityDefinition
Width-to-length ratioOutline width ÷ outline length, ×100. Conventionally the cephalic index; a ratio of 0.80 is an index of 80.
Left–right differenceDifference between two diagonals of the outline taken 30° either side of the vertical axis, ×100.

Dataset structure

One row per analysis window. All error figures are on the ×100 scale of the quantity concerned (cephalic-index points), not percentages of the value.

ColumnMeaning
windowWhich analysis window this row summarises
unique_photosPhotographs contributing to this row
ratio_rPearson correlation, automatic vs reference, width-to-length
ratio_mean_error / ratio_median_errorMean / median absolute error, width-to-length
ratio_biasMean signed error (automatic − reference); positive = automatic reads higher
ratio_within_{1,2,3,5}_pctShare of photographs within 1 / 2 / 3 / 5 points
lr_rPearson correlation, left–right difference
lr_mean_error / lr_median_errorMean / median absolute error, left–right difference
lr_within_{1,2}_pctShare within 1 / 2 points, left–right difference
three_class_agreement_pctAgreement on a three-way read of the ratio (thresholds 75 and 85)
large_lr_reference_nPhotographs where the reference recorded a clearly visible difference (≥ 3.5)
large_lr_both_n / large_lr_agreement_pctHow many of those the automatic measurement also recorded, and the share
large_lr_ours_only_nRecorded ≥ 3.5 automatically where the reference did not
deviation_over_5_n / deviation_over_5_pctPhotographs off by more than 5 points
mean_error_excluding_deviationsMean absolute error over the remaining photographs
detector_flag_rate_pctShare flagged by the automatic outline-quality check
detector_recall_pct / detector_precision_pctThat check's recall and precision on large deviations
large_deviation_clean_outline_pctLarge deviations whose outline nonetheless fits an ellipse well — failures the check cannot see
noteShort description of the window

The three windows

`primary_2024-05-07_onward` (n = 3,349) — the headline. One photograph per unique image, a single reference convention throughout, every reference value human-reviewed.

`excluded_before_2024-05-07` (n = 910) — published rather than discarded. Excluded from the headline because the reference measurement's own method was still being adjusted in that period: the reference median shifts by 3.0 points as an abrupt step around 2024-05-06/07, while the automatic measurement's own median is identical in both windows (81.9) and so is its segmentation quality (ellipse-residual median 0.0247 in both). The evidence points at the reference side rather than the measurement side — but the exclusion is a judgement call, so its numbers are here for anyone who wants to weigh it.

`all_windows` (n = 4,259) — both combined, published so the effect of the exclusion can be seen rather than taken on trust.

Uses

Suitable for: comparing photo-based anthropometric measurement against a human-reviewed reference; as a reference point for error rates in consumer measurement tools; as an example of a published first-party evaluation including failure rates.

Out-of-scope use

  • —Not a clinical validation. It compares two measurements of a photograph. It says nothing about health, development, or what any value means for an individual child.
  • —Not a population reference. This is a convenience sample, not balanced by age and not randomly drawn. The distribution describes the photographs measured, not the prevalence of any proportion among children.
  • —Reference values were reviewed by an experienced team, not by independent blinded assessors.
  • —Repeat-photograph consistency — the same head photographed twice, on different devices or in different light — is a separate question and is not addressed here.

Citation

This dataset is archived with a permanent DOI. Cite the DOI, not this page — it resolves for good, and it is the same two files, byte for byte.

doi:10.5281/zenodo.22706467

<https://doi.org/10.5281/zenodo.22706467>

Method: <https://www.headshapeai.com/methodology/> Full write-up, including two earlier revisions of this analysis and why they were wrong: <https://www.headshapeai.com/validation/> Benchmark page: <https://www.headshapeai.com/measurement-benchmark/>

Licence

CC BY 4.0 — reuse and redistribution permitted with attribution.