GPT-4 Can Predict Personality Test Answers Better Than You'd Expect
GPT-4 can create personality test questions that predict real-life outcomes like income, education, and well-being nearly as well as established psychology questionnaires. Researchers at Hebrew University of Jerusalem discovered this surprising capability by feeding the AI model text from two vastly different sources: a clinical psychiatry manual and an astrology guide. Both produced questions that correlated with participants' actual life circumstances, challenging assumptions about what makes a valid personality assessment.
How Did Researchers Test GPT-4's Personality Prediction Ability?
The team, led by Rotem Monsa, Aviv Zohar, and Shahar Arzy, conducted a rigorous experiment with 588 participants recruited through Prolific, a platform that pays people to participate in studies. Participants completed three personality assessments in random order: the Big Five Inventory (a standard 44-item psychology questionnaire), a 50-statement test that GPT-4 generated based on the DSM-5 (the diagnostic manual psychiatrists use), and a 60-statement test based on zodiac descriptions from "The Only Astrology Book You'll Ever Need." Researchers also collected detailed information about participants' lives, including income, education level, life satisfaction, anxiety and depression symptoms, employment status, and other behavioral measures.
The researchers deliberately chose these two source texts because both describe personality in rich detail, though only the DSM-5 has been refined through decades of clinical research. To guide GPT-4 toward questionnaire format, the prompts included example questions from the Big Five Inventory. Each statement the AI generated was limited to 10 words or fewer, and participants rated their agreement on a scale from 1 (strongly disagree) to 5 (strongly agree).
What Were the Surprising Findings About Astrology and Personality?
The results challenged conventional wisdom about personality assessment. On individual questions, the astrology-based test performed comparably to the DSM-5 and Big Five tests at predicting real-life outcomes. Most strikingly, on one measure,education level,the astrology questions actually outperformed the Big Five questions. However, when researchers combined answers into overall scores, the Big Five and DSM-based tests outperformed the astrology test on measures of mental health.
The astrology finding revealed something unexpected about how personality language works. When researchers grouped the astrology questions according to how participants actually answered them, ignoring the zodiac labels entirely, the four traditional elements (fire, water, earth, air) disappeared. Instead, four of the largest groups corresponded to Big Five personality traits: agreeableness, conscientiousness, extraversion, and openness.
"Astrological descriptions have been culturally refined over centuries and overlap heavily with everyday trait language, which is why the individual questions carry signal. The framework organizing them does not," explained Rotem Monsa.
Rotem Monsa, Researcher at Hebrew University of Jerusalem
What Does This Reveal About How GPT-4 Understands Human Behavior?
The study suggests that GPT-4 has absorbed fundamental patterns about how populations respond to personality questions, purely from reading vast amounts of human text. Before any participants answered questions, the researchers asked GPT-4 to estimate each question's average response on the 1-to-5 scale. For astrology questions, these estimates were off by an average of 0.27 points; for DSM-based questions, the model missed by 0.40 points. For every pair of questions, GPT-4 also estimated the correlation between answers, missing by an average of 0.17 on a scale from negative 1 to positive 1.
These prediction accuracies have practical implications. A tool that can estimate question quality before human testing could help researchers flag weak questions early, saving time and money. However, the researchers emphasized that this capability comes with important limitations and ethical boundaries.
Key Limitations and What Comes Next
- Language Dependency: The model was trained primarily on English text, so predictions would likely be less accurate in other languages. Lab members are currently testing this possibility across different linguistic contexts.
- Group-Level, Not Individual: The predictions apply to group averages rather than individual people. Predicting how a specific person will answer remains untested and separate from this research.
- Self-Reported Data: Every detail about participants' lives came from self-reports, which introduces potential bias compared to objective measures.
- Validation Requirements: Monsa stressed that researchers should never use an unvalidated test to diagnose patients or screen job candidates, regardless of who or what created it.
To demonstrate the limits of GPT-4's pattern recognition, the team repeated the procedure using a Bosch oven manual and a passage about Mordor from "The Lord of the Rings." GPT-4 still produced statements resembling personality questions, such as "Has well-arranged kitchen items" and "Rarely sees things positively." However, most of these statements were nonsensical, and researchers did not test them with participants. Useful questions emerged only from texts describing human behavior.
"The pull toward questionnaire format is strong even with nothing to build from," noted Monsa, highlighting how deeply the AI model has internalized the structure of personality assessment.
Rotem Monsa, Researcher at Hebrew University of Jerusalem
The team's next direction involves a more ambitious goal: reading personality directly from people's natural language without relying on questionnaires at all. Rather than asking people to answer structured questions, researchers want to use what GPT-4 has internally encoded about human personality to make predictions from conversation or writing samples.
"I wouldn't judge a questionnaire by who, or what, wrote it. I'd judge it by whether it has been validated," Monsa said, emphasizing that the source of questions matters far less than rigorous testing against real-world outcomes.
Rotem Monsa, Researcher at Hebrew University of Jerusalem
The full study was published in the journal iScience, and the team made every AI-generated question and all participants' responses available in a public archive for other researchers to examine. This transparency reflects a broader principle in AI research: that the validity of a tool depends on evidence, not on its creator.