Do AI models rate faces like humans?
We compared model judgements with human raters on a public dataset. Agreement is high on averageness and symmetry, lower on charm.
Why we asked
Our pipeline drafts every report with a vision model before a human reviews it. If model judgements drift from human perception, the human step has to catch it. We wanted to know where the drift is.
What we did
We took 500 frontal portraits from a public, consented research dataset with existing human attractiveness ratings, ran our assessment prompts, and compared the model's per-feature scores with the human averages.
What we found
- Averageness and symmetry: correlation above 0.8. The model sees what people see.
- Skin quality: strong agreement, slightly harsher than human raters on texture.
- Dimorphism: agreement is good for men, weaker for women, where the model overweights jaw width.
- "Charm" and expression: weak agreement. Humans reward warmth in a way models do not measure.
What it means for your report
The measured parts of a Faciem report are reliable. The parts about expression, uniqueness and charm are written by the reviewing analyst, not the model. That split is deliberate.