What happened
A new study published in the journal iScience revealed that OpenAI's GPT-4 model can generate personality surveys from different texts and predict, before the surveys are administered, the aggregate responses people will give to them. In the study conducted by Rotem Monsa, Aviv Zohar, and Shahar Arzy from the Hebrew University of Jerusalem and Hadassah Medical School, GPT-4 was asked to prepare two separate surveys—one based on the personality disorders section of the DSM-5, and the other based on an astrology book with no scientific validity.
The researchers note that while both sources address human personality in detail, they differ greatly in terms of scientific grounding. The surveys generated by GPT-4 were then administered to 600 adult participants, who also completed the Big Five Personality Inventory, a widely used tool in personality research.
What the findings show
In the DSM-5-based survey, traits that are related in real life, such as avoidance and dependency, were also found to move together in participants' responses; the internal consistency of this survey resembled the patterns found in the Big Five Personality Inventory. In the survey generated from the astrology book, however, zodiac sign and element groupings did not form consistent psychological dimensions, though some items still showed meaningful associations with outcomes such as depression, anxiety, and general well-being.
The most notable part of the study was GPT-4's predictions made before any data was collected from participants. The model was asked to forecast the average responses to each question on a five-point Likert scale, as well as the relationships between the questions.
- Correlation between predicted and actual average responses in the DSM-5-based survey: 0.71
- Correlation between predicted and actual average responses in the astrology-based survey: 0.85
- Performance in predicting relationships between questions: 0.74 and 0.69, respectively
The researchers emphasize that these figures are not accuracy rates, but correlation measurements showing the similarity between the model's predictions and actual participant data.
Why does it matter?
According to the researchers, since personality traits are reflected in everyday language, large language models (LLM) may have indirectly learned the general structure of human personality through their training data. This suggests that GPT-4 can function not just as a tool for generating survey questions, but also as an evaluator that predicts how people will collectively respond to the questions it creates.
The findings also indicate that large language models can extract measurable psychological signals even from texts with no scientific basis; this does not, however, validate astrology.
Limits and what comes next
The study's findings do not mean that GPT-4 knows what a specific individual thinks or how they would respond to a personality test. The model predicts the average tendencies of a group of 600 people and the statistical relationships between questions, rather than individual responses.
The researchers also note that the results do not provide evidence that these systems could replace clinical assessments or expert psychologists.
The real question: why was astrology predicted better?
The number that deserves the most attention is the higher correlation on the questionnaire with no scientific basis, not the one with it: 0.71 for DSM-5, 0.85 for astrology. That does not mean the model knows more about astrology; it most likely says something about what is actually being predicted.
Predicting item means in a survey is often predicting response style rather than content. People cluster toward the middle of a scale on statements that feel vague or unrelated to them, which makes a questionnaire with no meaningful latent structure easier to predict. The higher correlation may therefore be a sign of the instrument's weakness rather than the model's strength.
A second limit sits in the unit of analysis. The correlation is not computed across 600 people but across item means — so the effective sample is the number of questions, not the number of participants. A relationship fitted over a few dozen question averages carries a weaker claim than it sounds like it does.