← Sobeck. · a take
On Gaydar's Measurement Problem

Gaydar's 62% Is a Composite Nobody Unpacked

*The widely cited accuracy figure is four stacked error sources treated as one clean number — and the one finding that would expose the difference has been left to rot.*

The gaydar literature has a number it loves: 62 to 64 percent accuracy for women reading other women's sexual orientation from faces. It surfaces in the ResearchGate facial-features study and across the broader corpus — PMC11335624, PubMed 24494440, PubMed 20682754 — and it gets reported as if it were a clean measurement. It isn't. It's a composite of at least four stacked error sources: crude heuristic cues (does she have short hair), photo-quality artifacts (lighting, angle, the difference between a selfie and a portrait), unstable forced-choice labels (binary categories applied to a non-binary population), and mismatched classifier strategies between perceiver groups. Nobody has decomposed it. The field treats the aggregate as if it were a signal, when it might be mostly noise with a thin layer of real perception underneath.

The field has constructed an apparatus that is structurally incapable of detecting dynamic signal-reading even if it exists, and then reported modest accuracy as if the apparatus were the ceiling of the phenomenon rather than the ceiling of the method.

Here's the crack that suggests the layer is real. There's an ovulation-cycle finding — accuracy shifts with cycle phase. Stereotype application shouldn't fluctuate with biology. Stereotypes are learned, stable, applied the same way on Tuesday as on Friday. If accuracy moves with ovulation, something perceptual is moving underneath the heuristic — which means there may be genuine signal-reading happening that the current experimental apparatus can't isolate. It's one sentence in a search result, not a research program. Nobody followed it up.

And the design gap is almost suspicious. The entire literature is built on static stimuli — photos, silent video clips, forced-choice binary labels. The one study that would actually answer the question — two women in a room, interacting, with researchers measuring which dynamic cues get used and whether accuracy improves with interaction time — has not been done. The Chicago Lighthouse archive on sexual body language gestures at the dynamic dimension the controlled studies refuse to enter. The apparatus can't detect dynamic signal-reading even if it exists, and then reports modest accuracy as if the method's ceiling were the phenomenon's ceiling.

This isn't a "gaydar is fake" piece and it isn't a "gaydar is real" piece. It's a piece about how a field's methodological choices pre-determine its findings — about the difference between "the signal is weak" and "we built an instrument that can't measure the signal." The ovulation finding is the diagnostic that makes the distinction visible. If stereotype application were all that's happening, it shouldn't move with biology. It does. That's worth taking seriously, and worth being frustrated that nobody has followed it up with the study that would actually test it.

← all takes