The physics CALLUME · FIELD NOTE 14

Does color analysis work on dark skin? Yes. With two catches.

Yes, and the two ways it goes wrong on deep skin are both nameable. The physics are real: melanin masks the hemoglobin signal undertone rides on, cutting it by roughly 90 percent, which makes the vein test blind rather than wrong. The two failures that follow are arithmetic, not opinion: contrast measured as depth, and a saturation ceiling tied to the wrong number.

The physics: what an undertone test can and cannot see

Undertone is not a color you can see. It is a ratio. Melanin, hemoglobin, and carotenoid absorb in different bands, and the warm-cool call is read off the a* and b* axes of CIELAB. Melanin absorbs broadly across the visible spectrum, so it sits optically in front of the hemoglobin signature the temperature call depends on.

The number is about 90 percent. Peak hemoglobin signal drops roughly that much in dark skin, against about 70 percent in light skin, from a study of more than 15,000 spectra. It is physics, not a training-data problem. No amount of better software recovers a signal the pigment above it has already absorbed.

So the vein test is not wrong. It is blind here. It reads a subdermal cue melanin covers, which is why it degrades worst at the deep end. We cut it from our capture instructions at every skin depth, not only deep ones, because a test that works for some readers and quietly fails for others is not a test.

Second problem: hue angle goes unstable at low chroma. Deep skin sits closer to the neutral spine in Lab, and atan2(b*, a*) is mathematically unreliable near the achromatic origin. McLaren measured up to 35 degrees of hue error at low tristimulus ratios in 1980, and melanin-flattened skin sits exactly in that regime.

The reliability numbers say the same thing for everybody. Inter-rater agreement runs at kappa 0.37 for undertone and 0.45 for pigment, against 0.75 to 0.80 for depth (Weir et al., npj Digital Medicine, 2025). Undertone is the noisiest axis anyone measures. On deep skin it is noisier still, so the honest answer is a wider confidence interval on temperature at low skin chroma, with more weight on drape comparisons, which measure what a color does to a face rather than the absolute chroma of the skin.

And the standard depth metric cannot rescue it. ITA is arctan((L* − 50) / b*) times 180 over pi. It carries no a* term at all, so it cannot separate rosy from sallow at any lightness, and its category boundaries were originally fit on non-sun-exposed Caucasian subjects. It also destabilizes precisely where you need it: at an operating point of L* 30, b* values of 5, 7, and 9 return ITA readings of minus 76.0, minus 70.7, and minus 65.8 degrees. Roughly 5 degrees per unit of b*, and a one-unit wobble is well inside ordinary white-balance error.

Neutral skin hue itself rotates with melanin. The published range for neutral skin runs roughly 54 to 78 degrees, and global spectrophotometry puts real skin hue between 24.6 and 79.6 degrees, nearly double the range legacy tools inherited from European-only samples. A system that compares every face against one fixed pair of warm and cool anchors is using a ruler calibrated on somebody else. At deep lightness a fixed cool anchor can land almost exactly on true neutral while the warm anchor sits fifteen degrees off, and that asymmetric margin pulls deep skin cool for reasons that have nothing to do with the face in front of it.

Why almost every deep-skinned reader is called a Winter

Contrast is a subtraction, and on deep skin both numbers are dark. Most systems measure your contrast as the gap between your skin and your darkest feature, usually hair or eyes. On a light-skinned face, skin at L* 65 against an iris near L* 15 leaves a gap near 50: a genuinely high-contrast Winter, and no argument about it.

Now run the same subtraction on a deep-skinned face. Skin near L* 34, hair near L* 18, iris near L* 13. Skin-to-hair collapses to about 16 and the reading pins near the bottom of its range. The system prints low contrast. What it actually measured was the pigment corner, not the person.

The consequence is a mislabeled season, not a mismatched scarf. A deep-skinned Winter requires cool plus high chroma plus high contrast. Read as soft, she gets routed toward a Soft or Dark Autumn and loses the icy brights that are the entire point of her deck. It is the same arithmetic behind the field's oldest cliché, just wearing the opposite sign: measure depth and call it contrast, and every deep face becomes a Winter; measure the pigment corner and call it softness, and every deep face becomes a Soft Autumn. Either way a question about pigment got filed as a question about the person.

The same low reading is correct for a genuine Dark Autumn, whose blended medium-low contrast really does live there. That is what makes this a flaw rather than a bias correction. The ruler is not always wrong. It simply cannot tell the two cases apart.

The fix is to anchor contrast to the brightest feature the face actually owns. On most faces that is the sclera, and against a measured sclera a deep-skinned Winter presents the wide value gap she genuinely owns, in the same band a light-skinned Winter measures against her darkest feature. The contrast was never missing; the naive ruler was reading the wrong two points. One thing that fix must not do is presume the sclera is near-white: conjunctival melanosis, a benign darkening of the sclera, appears in roughly 92.5 percent of Black individuals (Singh et al. 1998, via Oh and Kanu 2020), so an honest version reads each session's own measured sclera rather than a constant.

And when it genuinely cannot tell, it should say so. If skin, hair, and eyes all sit at the dark end and the sclera reads dark too, there is no reliable value gap left to measure. The honest output there is a flagged low-confidence contrast reading settled by drape comparison, not a bare number printed as fact.

Read next: Warm vs cool →

The ceiling that cancels your brightest colors

The second failure is a saturation ceiling tied to the wrong number. Many systems cap how saturated a garment may be by reference to how saturated the face measures. That sounds principled and is quietly backwards, because facial chroma compresses as melanin rises. Read a perfectly ordinary deep-skin measurement against a scale built on light skin and it returns as extremely muted, which routes the reader into the Soft seasons and shaves the saturated colors off her palette on the way past.

Deep skin is not low-chroma skin. What compresses at the deep end is the measurement of the face, which is a fact about measuring faces, not a statement about wardrobes. Tie a garment ceiling to it naively and you demote exactly the colors that sing.

What is being measuredHexC*
Deep warm skin, a reference target#6D472Eabout 25
True Autumn, Pumpkin#FC7A1083.4
Bright Winter, Lemon Yellow#F6F54481.9
Dark Winter, Royal Purple#40038A78.6
Dark Autumn, Tomato Red#CB352273.8

Deck hexes are the shipped season colors, converted sRGB to Lab at D65. The skin hex is a reference target, not a captured measurement. The point of the table is the ratio: a face measuring around 25 has nothing to say about a garment measuring 83.

Cobalt, fuchsia, emerald, and marigold. Those colors on deep skin are a fashion cliché because they work. Any system that quietly capped them would reproduce the exact bias this field was supposed to beat, in a way no customer could ever audit from the outside.

The decision tree that fails black women

The bad tree fits on one line. Mechanically it is the contrast collapse above wearing a costume. TruHue names it exactly: “dark skin + dark hair + dark eyes = Dark Autumn or Dark Winter. Done.” Their own assessment is blunt: “for dark-skinned users the errors are massive,” because “the system can't see past your depth.” Winter demands cool plus high chroma plus high contrast. A deep, warm, muted person is an Autumn. A deep, warm, clear person is a Spring. A deep, cool, soft person is a Summer. Everyone deep is a Winter is what comes out when you measure depth and call it contrast.

The receipts do not sort by melanin. Analysts reading deep-skinned public figures put Jennifer Hudson and Solange Knowles on the clear, warm side, and Viola Davis and Serena Williams on the deep, muted side. Comparable skin depth, different families, and the separator is clarity. Those are third-party published analyst readings, cited here as commentary on public analyses rather than our measurements or any endorsement, and analysts do disagree on individual names, which is a property of a taxonomy with fuzzy boundaries and not a knock on the people reading it.

  • Style with DC: “Any skin tone can be in any season,” and cool tones “aren't exclusive to fair skin.”
  • Our own permanence rule, stated as law: depth is a dial reading, never a new season.

The variance data agrees, and it is not close. In one dataset, L* variance among Black participants measured 353 against 11 for White participants, and 89.4 percent of people have a perceptually indistinguishable counterpart in another ethnic group. Ethnicity is not a prior worth carrying. Measure the individual.

What depth actually changes: the value gap

Depth does not move the season. It moves which color does which job. The whole calculation is L* differences against skin. Here is the shipped True Autumn deck run against a deep skin operating point of L* 34.

Deck colorHexL*Gap from skinJob at this depth
Warm Indigo#3F4D6832.6about 1tonal match, base and structure only
Warm Brown#5B3B1C27.9about 6base garment, too thin at the collar
Dark Brown#432D1F20.7about 13the surviving dark anchor, your black as espresso
Golden Yellow#F7C40A81.4about 47a genuine contrast event
Warm Ivory#F8F6E896.7about 63the single biggest move this deck owns

Arithmetic computed from shipped deck hexes against a reasoned deep-skin operating point, not a captured session. Every color above stays in the palette. Nothing is demoted. What changes is placement.

The same math flips a very deep Bright Winter's power move outright. Against a skin band of L* 18 to 30, Pure White at L* 96.1 sits about 72 steps away from a band-center face, while the deck's near-blacks (Charcoal at L* 18.7, Midnight Blue at 16.3, Cool Navy at 10.7) sit inside her own band. Black still works, as an achromatic edge against chromatic skin, but it is no longer the largest lever. White is. That is a flip, not a gradient.

What we have not measured, stated plainly

Every constant behind this was tuned against one light-skinned reference. Skin #BB9689, L* 65.0, cool by 4.6 delta E. A deep-skin reference fixture is fully specified on paper, with target ranges and binary acceptance criteria. It has not been measured yet.

So read the corrections above as method, not as validated outcome. We can name the failure modes and say which ones we have addressed. We cannot yet show you a corrected deep-skin reading that has survived contact with a real face, and we are not going to pretend otherwise while the fixture sits unbuilt.

On accuracy: 75 to 85 percent, our own estimate. That is expected agreement with a careful human analyst. Season classification is genuinely hard for everyone. The best published four-season models plateau near 55 percent, twelve-subtype models run 30 to 60 percent, and no published inter-rater kappa for season assignment exists at all. We report stratified by depth rather than pooled, because pooled numbers hide the worst-served cell: Gender Shades found 34.7 percent error for darker-skinned women against 0.8 percent for lighter-skinned men, an intersection that vanishes the moment you average.

The at-home test that survives: the mirror ritual

Self-report tests fail. Comparison tests survive. The difference is what you are being asked to judge. Reading your own veins asks you to extract a pigment ratio from a cue melanin covers. Holding two objects at your face asks you to notice which one settles, which is a small draping session and works at any depth.

  • Gold against silver, one at a time at the jawline. Watch the shadow under your jaw and the skin under your eyes, not the metal.
  • Your cream shirt against your white one. The same two garments, the same light, back to back.
  • Overcast daylight only. Never store lighting. A 500 kelvin shift moves perceived undertone, and warm bulbs make nearly everyone read warm.
  • The bleed test, borrowed from working analysts: the right color bleeds into you, and you stop being able to see where the garment ends and the face begins. The wrong one draws a visible edge.

None of these is validated in the literature, and we say so. They are practitioner heuristics. What makes them better than the vein test is not proof, it is that they test the thing you actually care about: what a color does to your face. That mechanism, simultaneous contrast, has been documented since Chevreul in 1839, and it is the reason draping works at all.

One rule holds through all of it: the page is the script, never the instrument. No instruction anywhere in our documents tells you to hold a printed swatch against your wrist, at any skin depth, because screens and printers lie in ways nobody can audit from their own bathroom. Compare objects you own, in daylight, and let the measurement do the part measurement is good at.

Where this goes next. Building and measuring a deep-skin reference fixture is the next milestone. When it passes, the reasoned numbers on this page get replaced with measured ones. If it fails, the failure gets published the same way these two did: named and fixed in the open. If you want the measured version of any of this, that is what the product is for: a reading you can audit, on your own face, with the confidence bands printed.

Read next: Olive skin →

Questions

Yes, with a caveat that is physical rather than cultural. Depth, contrast, and clarity read cleanly. Undertone is the noisiest axis for everyone, at inter-rater kappa 0.37, and it is physically weaker at depth because melanin absorbs roughly 90 percent of the peak hemoglobin signal. The correct response is a wider confidence band on the temperature axis and more weight on side-by-side drape comparisons, not a different taxonomy.

Not a reliable self-report one. The vein test reads a subdermal cue that melanin covers, so it goes blind rather than wrong. What survives is comparison at the face in overcast daylight: gold against silver, your cream shirt against your white one, watching the shadow under your jaw rather than the object in your hand. That is a practitioner heuristic, and we label it as one.

Yes. Season is set by undertone, value contrast, and clarity, never by melanin, and every one of the four families contains deep-skinned members. Published analysts already place deep-skinned public figures on the clear warm side and the soft cool side alike, and the rule they state flat is that any skin tone can be in any season.

No, and what depth actually changes is more useful than what it does not. Depth moves the value arithmetic: which color does which job. At skin L* 34, a True Autumn's Warm Ivory sits about 63 lightness steps away and becomes the single biggest move in the deck, while Warm Indigo at a one-step gap becomes structure rather than statement. Same palette, different placement. The season itself does not move.

Check what it is actually measuring. If it asks about veins, or returns a season from lightness alone, it is reading depth. ITA, the standard depth metric, contains no red-green term at all, and it destabilizes exactly at the deep end: at L* 30, one unit of b* moves the angle about 5 degrees, which is inside ordinary white-balance error. A quiz built on that can rank depth. It cannot call temperature.

We expect 75 to 85 percent agreement with a careful human analyst, our own estimate rather than a measured figure, and no deep-skin fixture has been measured yet. The corrections on this page are engineering method with acceptance tests attached, not validated outcomes. We report accuracy stratified by skin depth rather than pooled, because pooled numbers hide the worst-served group.