
GPT-4 Predicts Aggregate Personality Responses from Text, But Not Individual Minds
Published by AINave Editorial • Reviewed by Ramit
A new study published in iScience shows that GPT-4 can predict how people will collectively respond to personality questions simply from the text used to write those questions. The model's predictions correlated with actual participant averages at 0.71 for questions derived from the DSM-5 and 0.85 for astrology-based items. But the researchers emphasize that this is about aggregate patterns, not individual-level measurement. For AI builders, this is a useful reminder that LLMs encode population-level response tendencies from language, which has implications for survey design, bias detection, and AI-assisted research tools.
How GPT-4 Predicted Questionnaire Responses
Researchers Rotem Monsa, Aviv Zohar, and Shahar Arzy from Hebrew University and Hadassah Medical School used GPT-4 to create two sets of personality questions. One set was drawn from the DSM-5, the standard clinical manual for personality disorders. The other came from an astrology textbook linking personality traits to zodiac signs. The team gave both questionnaires to 600 participants alongside the Big Five Inventory, the most validated personality measurement tool.
Before collecting any responses, GPT-4 estimated average scores for each item on a five-point Likert scale. Those predictions matched the actual participant averages strongly: a correlation of 0.71 for DSM-5 items and 0.85 for astrology items. The model also estimated relationships between items, yielding correlations of 0.74 for DSM-5 and 0.69 for astrology. The DSM-5-based questionnaire showed stronger internal consistency, meaning traits that appear together in real life also correlated in the responses.
The researchers also tested whether GPT-4 would generate personality-style questions from non-psychology sources. They fed it a Bosch oven manual and a landscape description from The Lord of the Rings. The model still produced questionnaire items resembling personality assessments, though some were semantically weak or narrow. This suggests GPT-4 has learned strong associations between language patterns, personality descriptions, and questionnaire structures -- even when the source material contains no explicit personality information.
What This Means for AI-Assisted Survey Design
LLMs are increasingly used to generate survey items, analyze text, or uncover biases. If GPT-4 can predict how people will answer questions based on language alone, that indicates these models have absorbed the statistical structure of human personality as a natural byproduct of training. For product teams building AI-assisted survey tools, this opens up possibilities: quickly generating plausible survey items, previewing response patterns, or detecting potential biases in existing questionnaires.
But the study also highlights a critical distinction. GPT-4 predicting aggregate averages is not the same as measuring individual traits. The researchers caution that predicting average responses does not establish that the resulting questionnaire accurately measures individual personality differences. A model that can forecast how a group will respond cannot be used to diagnose individuals. This matters when teams build AI tools for clinical assessment, hiring, or personalization.
Key Limitations to Keep in Mind
Several caveats matter for builders. The study used GPT-4 specifically; other models may behave differently. The conclusions are limited to aggregate prediction and may not generalize to other contexts or populations. When GPT-4 generated items from non-psychological sources, the outputs were sometimes semantically weak or poorly connected to the intended construct. The source material quality directly affects the quality of generated items: scientifically grounded texts like the DSM-5 produced better-structured questionnaires than astrology texts.
No claim was made about clinical utility or real-world diagnostic use. The study is a demonstration of language-driven response prediction, not a tool for individual assessment. For AI builders, the practical takeaway is clear: LLMs can surface population-level patterns from text, but using them to generate surveys or probes requires careful validation against established instruments, especially when the output influences decisions about real people.






















