Can People Fake Personality Tests? Yes: When Instructed
The fakeability research is unambiguous on one point: when people are explicitly instructed to present themselves as ideal candidates, they can and do shift their scores. The definitive meta-analysis pooled 51 studies of instructed faking and found that fakability did not vary by personality dimension, all five Big Five factors were equally fakable, and that distortion was largest on social desirability scales rather than on the trait scales themselves doi:10.1177/00131649921969802. Instructions to fake good produced smaller effects than instructions to fake bad, and within-subjects designs produced the more accurate estimates.
Two things are worth noting about that result. The first is what it does not say: it does not license a single headline number for "how much people inflate", because the estimate depends heavily on design, and the paper's own conclusion is that between-subjects designs distort the figure. The second is that instructed faking is a laboratory ceiling, not a description of hiring. The meta-analysis of real job applicants below found smaller mean differences than the instructed-faking studies did, and said so explicitly doi:10.1111/j.1468-2389.2006.00354.x.
This is not surprising. The social desirability bias, the tendency to answer questions in ways that present a favourable social image, is one of the best-documented phenomena in survey research. Personality tests are not immune to it. For a broader treatment of how assessment context shapes scores, see anonymity in personality assessment: why it matters.
Those two figures bracket the practical question. The first is what a conventional Likert-scale measure loses when the stakes go up, from a meta-analysis of 33 applicant samples doi:10.1111/j.1468-2389.2006.00354.x. The second is what a forced-choice measure loses under the same pressure, from a meta-analysis of studies comparing forced-choice scores across low-stakes and high-stakes conditions doi:10.1037/apl0000414. Both sections below unpack them.
How Much Faking Actually Happens in Real Hiring Contexts
Here is where the picture becomes more nuanced. The laboratory evidence for faking under instruction is robust. The evidence for spontaneous, uninstructed faking in actual hiring contexts is weaker.
The relevant meta-analysis pooled 33 studies comparing job applicants with non-applicants on the same personality measures. Applicants scored higher on conscientiousness (d = .45), emotional stability (d = .44), openness (d = .13) and extraversion (d = .11); agreeableness was not among the dimensions showing a significant applicant advantage, and the pattern shifted by job type, with applicants distorting most on the dimensions most obviously relevant to the job they wanted doi:10.1111/j.1468-2389.2006.00354.x. So the differences are real, concentrated on two of the five dimensions, and smaller than instructed faking implies.
A large retest study pushes further in the same direction. Among 5,266 rejected applicants who reapplied for the same job six months later and retook the same five-factor measure, 5.2% or fewer improved their scores on any scale, scores were as likely to move down as up, and only three people improved across all five scales beyond a 95% confidence threshold doi:10.1037/0021-9010.92.5.1270. People who had every reason to game the test on the second attempt overwhelmingly did not.
Some researchers argue the remaining gap is mostly attributable to self-deception rather than deliberate distortion: applicants genuinely believe they possess the traits they claim, because motivated reasoning leads them to inflate their self-assessments in positive directions. This is a different phenomenon from cynical faking, and it has different implications for what to do about it. It also connects to the broader evidence on why self-assessment alone produces an incomplete picture: we are systematically unreliable narrators of our own personalities.
Does Faking Actually Damage Predictive Validity?
This is the key empirical question, and the answer is genuinely contested. There are three main positions in the literature:
Position 1: Faking doesn't matter much. The meta-analytic case here is direct. Social desirability scales were found not to predict job performance, school success, task performance or counterproductive behaviours, and social desirability did not function as a useful suppressor or mediator for job performance; removing its effects from the Big Five left their validity for predicting workplace performance intact doi:10.1037/0021-9010.81.6.660. If fakers end up performing just as well on the job, the faking has not corrupted the predictive signal.
Position 2: Faking reduces validity. Other research argues that when high scorers are a mixture of genuinely high-trait individuals and skilled fakers, the predictive relationship with performance is attenuated. The fakers inflate the distribution without having the underlying trait.
Position 3: Faking selects for impression management. A third perspective notes that the ability to fake a personality test successfully may itself be a meaningful signal, specifically, it correlates with social intelligence and the motivation to impress. People who fake effectively may actually perform well in roles where impression management matters. This does not make faking unproblematic, but it complicates the simple validity-corruption story.
The current scientific consensus leans toward Position 1 for low-stakes, non-selection uses, and toward Position 2 for high-stakes selection contexts where the incentive to fake is highest. For what "valid" actually means in personality measurement, see what is reliability and validity in personality testing?.
Forced-Choice Formats: A Partial Solution to Test Faking
One technical response to faking is the forced-choice format, in which respondents choose between two or more equally socially desirable options (e.g., "I tend to be very organised" vs. "I tend to be very sociable"). By removing the easy option of endorsing all positive descriptors, forced-choice formats reduce the opportunity for response inflation.
The evidence is good, and it is quantified. A meta-analysis comparing forced-choice scores between low-stakes and high-stakes conditions put the overall score inflation for forced-choice measures at d = 0.06, far below the equivalent figure for single-statement measures across most facets, and identified which design choices matter: statements balanced for social desirability, normative rather than purely ipsative scoring, and multidimensional blocks doi:10.1037/apl0000414. A second meta-analysis reached a compatible conclusion, that forced-choice inventories resist faking, that the effect is larger in experimental settings than in real selection, and that quasi-ipsative formats resist best, with conscientiousness inflation of δ = 0.49 under a quasi-ipsative format against δ = 1.27 under a fully ipsative one doi:10.3389/fpsyg.2021.732241.
The cost is real too. Forced-choice items are harder to take, more cognitively demanding, and less well-validated in cross-cultural contexts. Purely ipsative scoring makes scores relative to each other rather than absolute, which complicates interpretation, and the second meta-analysis above shows the format choice inside "forced-choice" matters as much as the decision to use it. For a detailed treatment of the tradeoffs, see forced-choice personality assessment: more honest data.
| Test format | Susceptibility to faking | What to do |
|---|---|---|
| Standard Likert-scale Big Five | High, easy to endorse positive items | Use for development; monitor if used in selection |
| Forced-choice (e.g., Thurstonian IRT) | Low overall (d = 0.06), but depends on how it is built | Better for selection contexts; validate locally |
| Situational judgement tests (SJTs) | Moderate, less transparent but still fakeable | Combine with personality data for better coverage |
| Peer/multi-rater assessment (e.g., Cèrcol Witnesses) | Low, others are harder to game | Most resistant to intentional faking |
| Anonymous self-report (no high-stakes context) | Low, reduced motivation to distort | Best for development; the ideal context for honest data |
See also: social desirability bias in personality tests.
Why the Real Problem Is Using Personality Tests for Selection
Most of the faking debate is actually a proxy debate about whether personality tests should be used for hiring selection at all. If you use personality data for development, to help people understand themselves and work better with others, faking is largely self-defeating. If you fake your way to a development report that doesn't reflect how you actually behave, you will simply get feedback that is less useful to you. The cost falls on the person who faked, not on the organisation.
If you use personality data as a selection hurdle, as a gate that determines whether you proceed in a hiring process, the incentive to fake becomes significant, and the consequences of faking fall on the organisation (in the form of mis-selected hires) and on honest candidates who are disadvantaged relative to skilled fakers. The ethical and legal dimensions of this are examined in personality testing in hiring: what is legal and what is ethical.
This is the most important practical insight from the faking literature: the fakeability problem is downstream of the misuse problem. Personality assessment was designed for development and self-understanding. When it is repurposed as a selection screen, it develops fakeability vulnerabilities that it was never designed to resist.
Why Anonymous, Developmental Assessment Resists Faking
One practical response to faking concerns is to collect personality data anonymously: that is, in contexts where the data will not be seen by hiring managers, and where no individual's scores are linked to employment decisions. This is exactly the context in which personality assessment generates its most valid and useful signal.
Anonymous assessments eliminate the primary motivation for faking. They produce data that is more accurate and more useful for development. They also eliminate the legal and ethical risks associated with using personality data in selection.
Cèrcol uses a multi-rater design in which Witnesses provide an external view of a person's behaviour. This adds a layer of information that is substantially harder to fake than self-report: while you can inflate your own scores on Conscientiousness, you cannot easily manipulate how five colleagues who work with you every day describe your typical behaviour. The science behind the IPIP: International Personality Item Pool that underpins the Witness instrument gives it validated, open-source foundations that commercial tools frequently lack.
The faking literature ultimately teaches us less about how to make personality tests more tamper-resistant and more about where they genuinely belong: in development, in team dialogue, in self-understanding. That is where they deliver their most honest and most useful signal.
How Cèrcol addresses faking
Cèrcol's design responds directly to the faking problem at every level. The Witness peer assessment means that your personality profile is not constructed from your own self-presentation: it is built from how the people who work with you actually describe your behaviour. Witnesses answer forced-choice adjective pairs anonymously, which removes both the ability to fake (no obviously "correct" answer) and the incentive (no stakes attached to their individual responses). The self-report component is taken in a low-stakes, individually owned context, you own your data and decide who sees it, which research shows substantially reduces impression management. You can explore the full instrument at /instruments and read about the underlying science at /science. The assessment is free at cercol.team.
Sources
- Viswesvaran, C., & Ones, D. S. (1999). Meta-analyses of fakability estimates: Implications for personality measurement. Educational and Psychological Measurement. https://doi.org/10.1177/00131649921969802
- Birkeland, S. A., Manson, T. M., Kisamore, J. L., Brannick, M. T., & Smith, M. A. (2006). A meta-analytic investigation of job applicant faking on personality measures. International Journal of Selection and Assessment. https://doi.org/10.1111/j.1468-2389.2006.00354.x
- Hogan, J., Barrett, P., & Hogan, R. (2007). Personality measurement, faking, and employment selection. Journal of Applied Psychology. https://doi.org/10.1037/0021-9010.92.5.1270
- Ones, D. S., Viswesvaran, C., & Reiss, A. D. (1996). Role of social desirability in personality testing for personnel selection: The red herring. Journal of Applied Psychology. https://doi.org/10.1037/0021-9010.81.6.660
- Cao, M., & Drasgow, F. (2019). Does forcing reduce faking? A meta-analytic review of forced-choice personality measures in high-stakes situations. Journal of Applied Psychology. https://doi.org/10.1037/apl0000414
- Martínez, A., & Salgado, J. F. (2021). A meta-analysis of the faking resistance of forced-choice personality inventories. Frontiers in Psychology. https://doi.org/10.3389/fpsyg.2021.732241
- Wikipedia: Social desirability bias
- IPIP: International Personality Item Pool: https://ipip.ori.org
Further reading
- Social desirability bias in personality tests
- Forced-choice personality assessment: getting more honest data
- Why self-assessment alone isn't enough: the case for peer personality feedback
- Anonymity in personality assessment: why it matters
- Personality testing in hiring: what is legal and what is ethical?
- What is reliability and validity in personality testing?