AI Bias in Facial Judgments: Four Models Infer Character
psychologytoday.com

AI Bias in Facial Judgments: Four Models Infer Character

Tech News
3 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRA report on a Harvard study describes four multimodal AI models making character judgments from facial images. The strongest comparison with human bias was specific to competence, not trustworthiness.

AI bias in facial appearance judgments showed up across four models in a Harvard study summarized by Psychology Today: the systems inferred traits such as trustworthiness and competence from face images. The report says none of the models responded neutrally or refused the biased prompts tested, despite industry efforts to reduce undesirable behavior. The study covered nearly 8,000 trials across 13 experiments.

What the tests found, and what they did not

The team tested GPT-4o and GPT-5 from OpenAI, Claude Sonnet 4.5 from Anthropic, and Gemini 3 Flash Preview from Google. The experiments used two-dimensional face images to probe judgments of trustworthiness and competence. The report gives no model-by-model scores, so it supports a shared finding across these systems, not a ranking of which was most biased. The four models and experiment scope are described in the article.

One result deserves precise wording: the models showed higher-magnitude bias than humans when judging competence from faces. The report does not make the same stronger-than-human comparison for trustworthiness, nor does it provide a numerical effect size. That distinction matters because “more biased than people” is broader than the reported result supports. The comparison is specific to competence judgments.

Why this matters for selection tools

A face-based judgment can turn appearance into a proxy for a quality the system cannot establish from an image. The researchers warned against using language models in selection contexts; the report also says the tested prompts included consequential judgments involving criminality and worthiness of financial investment. These examples indicate the kind of inference at issue, but the experiments did not measure hiring outcomes or demonstrate a specific discriminatory decision in a real hiring process. The report describes the researchers’ warning and the consequential prompts.

That makes the finding relevant to any workflow that invites a model to assess a person from a photo. A score or summary might look like a judgment about capability or character while resting on facial appearance instead. The risk follows from the behavior tested, not from evidence that the study evaluated a deployed hiring system.

Alignment goals are not a guarantee

The article notes that models such as GPT-4o are trained with methods intended to avoid socially undesirable bias and align outputs with human goals. Yet the models in these experiments still made face-based character inferences and did not refuse the tested prompts. The researchers suggest such inferences may emerge from models trained on human knowledge even without deliberate instruction to make them. The article describes both the alignment goal and the observed responses.

This does not show that alignment training cannot reduce facial bias: the reported study did not compare training interventions. It does show why a stated safety goal and neutral behavior in a specific multimodal task are separate things to establish. Until that distinction is tested directly, appearance-based judgments remain a poor substitute for evidence about a person’s abilities or character.

FAQs

The report says the tested models inferred trustworthiness and competence from facial images, and none responded neutrally or refused the biased prompts tested. The experiments tested both traits.

Sources

Latest Tech News