GPT-4 Predicts Aggregate Personality Responses from Text, But Not Individual Minds
techradar.com

GPT-4 Predicts Aggregate Personality Responses from Text, But Not Individual Minds

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRGPT-4 can forecast average personality responses from text, as shown in a new iScience study. The model predicted actual participant averages with correlations of 0.71 (DSM-5) and 0.85 (astrology), but this does not imply individual-level measurement. AI builders should understand the distinction for survey design and bias considerations.

A new study published in iScience shows that GPT-4 can predict how people will collectively respond to personality questions simply from the text used to write those questions. The model's predictions correlated with actual participant averages at 0.71 for questions derived from the DSM-5 and 0.85 for astrology-based items. But the researchers emphasize that this is about aggregate patterns, not individual-level measurement. For AI builders, this is a useful reminder that LLMs encode population-level response tendencies from language, which has implications for survey design, bias detection, and AI-assisted research tools.

How GPT-4 Predicted Questionnaire Responses

Researchers Rotem Monsa, Aviv Zohar, and Shahar Arzy from Hebrew University and Hadassah Medical School used GPT-4 to create two sets of personality questions. One set was drawn from the DSM-5, the standard clinical manual for personality disorders. The other came from an astrology textbook linking personality traits to zodiac signs. The team gave both questionnaires to 600 participants alongside the Big Five Inventory, the most validated personality measurement tool.

Before collecting any responses, GPT-4 estimated average scores for each item on a five-point Likert scale. Those predictions matched the actual participant averages strongly: a correlation of 0.71 for DSM-5 items and 0.85 for astrology items. The model also estimated relationships between items, yielding correlations of 0.74 for DSM-5 and 0.69 for astrology. The DSM-5-based questionnaire showed stronger internal consistency, meaning traits that appear together in real life also correlated in the responses.

The researchers also tested whether GPT-4 would generate personality-style questions from non-psychology sources. They fed it a Bosch oven manual and a landscape description from The Lord of the Rings. The model still produced questionnaire items resembling personality assessments, though some were semantically weak or narrow. This suggests GPT-4 has learned strong associations between language patterns, personality descriptions, and questionnaire structures -- even when the source material contains no explicit personality information.

What This Means for AI-Assisted Survey Design

LLMs are increasingly used to generate survey items, analyze text, or uncover biases. If GPT-4 can predict how people will answer questions based on language alone, that indicates these models have absorbed the statistical structure of human personality as a natural byproduct of training. For product teams building AI-assisted survey tools, this opens up possibilities: quickly generating plausible survey items, previewing response patterns, or detecting potential biases in existing questionnaires.

But the study also highlights a critical distinction. GPT-4 predicting aggregate averages is not the same as measuring individual traits. The researchers caution that predicting average responses does not establish that the resulting questionnaire accurately measures individual personality differences. A model that can forecast how a group will respond cannot be used to diagnose individuals. This matters when teams build AI tools for clinical assessment, hiring, or personalization.

Key Limitations to Keep in Mind

Several caveats matter for builders. The study used GPT-4 specifically; other models may behave differently. The conclusions are limited to aggregate prediction and may not generalize to other contexts or populations. When GPT-4 generated items from non-psychological sources, the outputs were sometimes semantically weak or poorly connected to the intended construct. The source material quality directly affects the quality of generated items: scientifically grounded texts like the DSM-5 produced better-structured questionnaires than astrology texts.

No claim was made about clinical utility or real-world diagnostic use. The study is a demonstration of language-driven response prediction, not a tool for individual assessment. For AI builders, the practical takeaway is clear: LLMs can surface population-level patterns from text, but using them to generate surveys or probes requires careful validation against established instruments, especially when the output influences decisions about real people.

FAQs

In the cited study, GPT-4's estimated average scores for questionnaire items correlated with actual participant averages at 0.71 for DSM-5-based items and 0.85 for astrology-based items, indicating strong aggregate predictability. However, the researchers caution that this does not imply accuracy at the individual level.

Sources

Latest Tech News