The evidence CALLUME · FIELD NOTE 05

Is color analysis accurate? Every published number we could find.

Nobody can currently prove it either way: in roughly forty-six years of practice, no study has ever published how often two analysts give the same person the same season. Numbers do exist for the ingredients — trained raters agree on undertone at only κ = 0.37 (Weir et al., 2025) — and Callume states its own expectation as an estimate: roughly 75 to 85 percent agreement with an in-person analyst.

Accurate against what?

The first problem with the question is that there is nothing objective to be accurate against. No instrument returns “warm” or “cool”; no lab test says Autumn. A season is a judgment built on top of measurements — so in practice, “accurate” can only mean agreement: between two analysts, between two tools, between your Tuesday result and your Friday one.

Researchers score that kind of agreement with kappa (κ), a statistic corrected for luck: 0 means agreement no better than chance, 1 means perfect. As a rough convention, values above 0.6 read as substantial and values around 0.4 as fair. Hold onto that scale — every number below uses it or a close cousin.

Every published number we could find

What was measuredResultSource
Trained raters agreeing on skin undertone, in person (Pantone SkinTone guide)κ = 0.37 — “fair”Weir et al., npj Digital Medicine, 2025
The same raters agreeing on pigmentκ = 0.45Weir et al., 2025
The same raters agreeing on depth (Monk scale)κ = 0.75 in person; up to 0.80 by body siteWeir et al., 2025
A second team, different cohort, agreeing on Monk depthICC = 0.64Cu et al., npj Digital Medicine, 2025
A colorimeter re-measuring the same skin — an instrument, no humansICC = 0.98Weir et al., 2025
Novice raters handed a strict protocol (fixed crop, enforced white balance)ICC = 0.93–0.98PMC3778111
Natural clusters in 100,000+ skin measurements matching expert seasons36% (chance = 25%)Gaston
Best machine-learning classifiers on the purpose-built season dataset~55% (4 seasons); 30–60% (12 sub-seasons)Deep Armocromia benchmark
Generic color-wheel harmony predicting skin-harmony judgments~53% — barely above chanceSirisayan (Chulalongkorn)
Two analysts assigning the same person the same seasonNever published

Two caveats belong beside that table, not under it. The clustering figure comes from a study whose cluster-counting method — the “elbow” heuristic — is itself ambiguous, so treat the 36 percent as suggestive rather than surgical. And Deep Armocromia, the field's best public benchmark, is roughly 86 percent lighter-skinned and has published no inter-annotator agreement of its own: the benchmark for season accuracy carries the same measurement gap as the field it benchmarks.

Read the table top to bottom and a pattern falls out. Instruments agree with themselves almost perfectly. Humans given strict protocols agree well. Humans judging depth agree substantially. And agreement collapses exactly where season systems need it most — undertone, the warm-versus-cool call, lands at 0.37 among trained professionals holding a reference guide in daylight.

The number nobody has published

As of August 2026, no inter-rater agreement figure for season assignment has ever been published — not a kappa, not a percentage, not once. Our research survey records it as a confirmed gap, and at least one competitor's own technical white paper concedes the neighboring point: the draping test has no peer-reviewed validation, and the twelve-season taxonomy evolved through practitioner experience rather than measurement.

That sentence covers us too. There is no published benchmark to be accurate against, so the honest options are exactly two: publish the study, or state your expectation as an estimate and label it one. We chose the second, and we are working toward the first.

Even the founder of the mass-market system was candid about this. Carole Jackson, whose 1980 Color Me Beautiful put four seasons into millions of homes, called the seasons “a convenient artifice.” The label is a communication device stretched over real, continuous measurements — useful exactly as long as nobody mistakes it for a law of nature.

Where the misses concentrate

The weakest axis is undertone, and the physics says why. Melanin optically masks the hemoglobin signal that undertone judgments read on — in deep skin the peak hemoglobin signal drops by around 90 percent, against roughly 70 in light skin (spectral studies of 15,000+ measurements) — so the axis is genuinely harder to measure at low lightness, not merely under-practiced. Every folk shortcut, from the vein test to the jewelry test, degrades on exactly the same skin.

Which is why one pooled accuracy number would mislead even if it existed. The pattern is documented across face-analysis systems: the Gender Shades audit found commercial classifiers erring on darker-skinned women at 34.7 percent against 0.8 percent for lighter-skinned men — a forty-fold gap that a single average hides completely. An honest accuracy claim in this field has to say accurate for whom, and a reading should print its uncertainty per person, not per marketing page.

Callume's own number, stated as what it is

We expect photo-based color analysis to land around 75 to 85 percent agreement with an in-person analyst — our own estimate, since no published figure exists — and Callume caps its stated confidence at 95 percent on purpose. Every axis in a reading carries an error bar, every facial measurement a reliability tier, and the verdict names the sister season that nearly won. The standing policy lives on our honesty page, linked below.

The estimate is deliberately humble about hardware, too. In dentistry — a far better-funded color-matching field — instrument-assisted matching beats trained human judgment by only 0.02 to 0.15 κ. Calibration removes the lighting variable from your photos; it does not manufacture certainty. Anyone telling you a gadget fixed reliability is overselling the gadget.

What would settle it

The missing study is not exotic. Take a stratified panel of clients across the full range of skin depths; have analysts from several lineages — Sci\ART, House of Colour, the Korean schools — and the leading tools each assign a season independently; publish the agreement as a kappa against a pre-registered bar (0.667 is a conventional pass line). Forty-six years in, nobody has run it. We've scoped it, our blind drape-voting data is structured to feed it, and if it ships, this page is where the field's first published figure will live.

This page is maintained as a living survey. If a published reliability figure for color analysis exists that isn't listed here, tell us and we'll add it — the point of this page is that the list is complete.

Read next: Why results disagree →

Questions

Is color analysis scientifically proven?

No — and nobody should tell you otherwise. Draping has no peer-reviewed validation, and no study has ever published analyst agreement on season assignment. What is solid is the underlying colorimetry: skin measured in CIELAB by instrument is near-perfectly repeatable (ICC 0.98). The science supports measuring your coloring; the twelve labels on top are a useful convention, not a discovered fact.

How accurate are color analysis apps?

Strictly: unknown. There is no published benchmark for season assignment, so no app's accuracy percentage can be verified — including ours, which is why Callume publishes an estimate (we expect roughly 75 to 85 percent agreement with an in-person analyst) instead of a claim. What is documented is that uncalibrated photo reads move with your lighting.

What is the most accurate way to get my colors done?

By the numbers available: an in-person analyst draping under controlled light remains the reference standard, and a calibrated photo read — one that corrects your lighting against a known reference in the frame — is the closest a tool gets. An uncalibrated app read is the least stable of the three. Whichever you choose, weight results that repeat over results that merely sound sure.