How We Evaluated Free Personality Tests: The Validity Criteria
The first figure comes from the meta-analysis that established personality as a legitimate predictor in personnel selection. Barrick and Mount reported an estimated true score correlation of ρ = .22 between Conscientiousness and job performance as a mean across criterion types, and the same ρ = .22 as the mean across five occupational groups doi:10.1111/j.1744-6570.1991.tb00688.x. Note what that value is and is not: it is a corrected correlation for one trait, not for the Big Five as a set, and the uncorrected observed mean in the same table is r = .13.
The second figure is the reliability benchmark any instrument should be judged against. Gnambs meta-analysed 682 test-retest correlations from 74 samples (total N = 14,923) collected within intervals of up to two months, and reported a median dependability estimate of ρtt = .816 across the five traits doi:10.1016/j.jrp.2014.06.003.
The third card is a fact about this site rather than a research finding: First Quarter Cèrcol administers the 60-item IPIP-NEO-60 and returns 30 facet scores at no cost. The instrument itself is documented in Maples-Keller et al. (2019).
Every tool in this ranking is assessed on five dimensions:
- Peer-reviewed validation: Has the instrument been independently validated in published, peer-reviewed research, not just by the vendor?
- Public domain / open source: Are the items and scoring algorithms publicly inspectable, or is the methodology proprietary?
- Test-retest reliability: Do scores remain stable over weeks and months, as personality theory predicts they should?
- Predictive validity: Do scores predict outcomes that matter, job performance, team effectiveness, wellbeing?
- Free to use: Can a team use it without per-seat licensing costs?
For a primer on what these criteria mean statistically, see what is reliability and validity in personality testing.
The Best Free Personality Tests for Teams in 2026, Ranked by Science
1. IPIP-based assessments: strongest evidence, free, open
What it measures: All five Big Five dimensions (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) as continuous scores, with item-level transparency.
Scientific basis: The International Personality Item Pool is a public-domain library of personality items developed by Lewis Goldberg and now hosted by the Oregon Research Institute. IPIP scales were built as public-domain proxies for commercial inventories, and the convergence is documented: Johnson reports that the IPIP-NEO scales correlated on average r = .73 with the NEO PI-R scales they were modelled on, rising to r = .94 when corrected for attenuation due to scale unreliability doi:10.1016/j.jrp.2014.05.003. The five-factor structure itself replicates across languages: McCrae and Costa compared six translations of the NEO PI-R against the American factor structure (N = 7,134, five distinct language families) and concluded that the data strongly suggest personality trait structure is universal doi:10.1037/0003-066X.52.5.509. To understand how these items became the scientific standard, see what is the IPIP and why does it matter and history of the Big Five from Allport to Goldberg.
Test-retest reliability: High. The meta-analytic median dependability estimate for Big Five scales is ρtt = .816 over intervals of up to two months, with transient error accounting for roughly 10 per cent of observed variance in trait scores doi:10.1016/j.jrp.2014.06.003.
Predictive validity for job performance: The strongest personality evidence in personnel selection concerns Conscientiousness. Barrick and Mount concluded that "one dimension of personality, Conscientiousness, showed consistent relations with all job performance criteria for all occupational groups", at ρ = .22 doi:10.1111/j.1744-6570.1991.tb00688.x. In the same analysis, Extraversion was a valid predictor for managers and sales roles (ρ = .18 and .15), and both Openness and Extraversion predicted training proficiency (ρ = .25 and .26).
Free to use: Yes. The IPIP states that its items and scales are in the public domain and can be copied, edited, translated or used for any purpose without permission or fee.
Team-ready: With appropriate presentation and interpretation, yes, but raw IPIP administration requires some setup.
2. Cèrcol: IPIP instruments plus Witness peer assessment
What it measures: All five Big Five dimensions (as Presence, Bond, Vision, Discipline, Depth) via self-report, plus structured Witness assessments from peers chosen by each participant.
Scientific basis: Built directly on IPIP items. First Quarter Cèrcol is the 60-item IPIP-NEO-60, whose scores showed good reliability and convergent validity with the NEO PI-R and the IPIP-NEO-300, and matched the NEO-FFI's nomological network across a wide range of external variables doi:10.1080/00223891.2017.1381968. Full Moon Cèrcol uses the 120-item IPIP-NEO-120, whose psychometric properties Johnson found to "compare favorably to the properties of the longer form" doi:10.1016/j.jrp.2014.05.003. New Moon Cèrcol is a ten-item screener based on the TIPI, which its authors describe as somewhat inferior to standard multi-item instruments while still reaching adequate levels of convergence and test-retest reliability doi:10.1016/S0092-6566(03)00046-1. Use it as an orientation, not as a profile.
The Witness peer assessment layer adds a second perspective that self-report alone cannot capture. This is not a soft claim: in a meta-analytic integration covering 44,178 targets across 263 samples, Connelly and Ones found that when the criterion was academic achievement or job performance, other-ratings yielded predictive validities substantially greater than, and incremental to, self-ratings doi:10.1037/a0021212. See why self-assessment alone isn't enough: peer personality feedback for the wider evidence base.
Test-retest reliability: Inherits the stability of the IPIP Big Five baseline described above.
Predictive validity: The self-report instruments carry the validity of the IPIP-NEO scales they implement. The Witness layer is a peer-rating design of the kind Connelly and Ones found to add incremental prediction, though Cèrcol has not yet published criterion validity data of its own.
Free to use: New Moon Cèrcol and First Quarter Cèrcol are free. Full Moon Cèrcol, which adds the 120-item instrument and the Witness assessments, is a one-time paid purchase, and is free to new accounts while the open beta lasts.
Team-ready: Yes: specifically designed for team use. Results include a team-level view showing the distribution of traits across the group. See what the Cèrcol Witness instrument measures for the full instrument description.
3. Open Source Psychometrics Project
What it measures: Multiple instruments available for free, including a Big Five test built from public-domain International Personality Item Pool scales.
Scientific basis: The Open Source Psychometrics Project hosts well-documented personality instruments and publishes the anonymised raw response data behind them. Its IPIP Big Five Factor Markers dataset alone holds 1,015,342 responses, and the site states that its datasets have been used in more than 25 published papers. That level of transparency is unusual and is the main reason the project ranks here.
Test-retest reliability: Varies by instrument; the IPIP-based measures inherit the Big Five dependability estimates above.
Predictive validity: Carried by the underlying IPIP scales rather than established for the site itself.
Free to use: Yes, fully free.
Team-ready: Results are individual; aggregating into team profiles requires additional work. Suitable for teams comfortable with some self-service interpretation.
4. 16Personalities: popular, limited independent validation
What it measures: Five scales, presented by the vendor as Energy, Mind, Nature, Tactics and Identity, where Identity is the Assertive/Turbulent axis. The first four are reported using MBTI-style letters, producing 16 named types.
Scientific basis: The vendor states that its NERIS Type Explorer keeps the four-letter acronym format but bases the model on five-factor traits rather than on Jungian cognitive functions. The mapping between the letters and the five factors has been studied directly for the MBTI: McCrae and Costa found "no support for the view that the MBTI measures truly dichotomous preferences or qualitatively distinct types", while the four indices did measure aspects of four of the five major dimensions of normal personality doi:10.1111/j.1467-6494.1989.tb00759.x. That is the core problem with any type output: the underlying variation is continuous, and sorting people into boxes throws away the part of the score that carries the information. No independent peer-reviewed validation of the NERIS instrument itself was found for this review. For a thorough breakdown, see 16Personalities vs Big Five: the viral test that gets it half right.
Test-retest reliability: Not established in independent published research that we could locate for this instrument.
Predictive validity: Likewise not established in independent research.
Free to use: Yes. The vendor states the test is free and requires no registration.
Team-ready: Widely used as an icebreaker and conversation starter. Should not be used for high-stakes decisions about roles or development.
5. MBTI: commercial, and disowned for selection by its own publisher
What it measures: Four dichotomies yielding a four-letter type. Form M, the standard form, has 93 forced-choice items.
Scientific basis: Jung's theory of types, operationalised by Myers and Briggs. Two reviews by Pittenger concluded that the available literature offers insufficient evidence to support the tenets of and claims about the utility of the test doi:10.3102/00346543063004467, and that the instrument "may not yet be able to support the claims its promoters make" doi:10.1037/1065-9293.57.3.210. The most recent synthesis is blunter about what is missing: aggregating 193 studies published between 1999 and 2024, Erford and colleagues found internal consistency of .845 to .921 across subscales and total scores, but reported that "structural validity and test-retest studies were absent from the 25-year literature sampling" doi:10.1002/jcad.70006. An instrument can be internally consistent and still fail to show that its four-box structure is real or that a person's type holds still.
Free to use: No. Form M is sold through licensed distributors, in Canada at USD 382 per pack of ten for the paper version, and is restricted to qualified purchasers.
Team-ready: For discussion, yes. For selection, the publisher itself says no: "The MBTI assessment should never be used in recruitment or selection because it does not measure a person's skills or abilities."
6. DISC: proprietary, limited independent validation
What it measures: Four behavioural styles (Dominance, Influence, Steadiness, Conscientiousness), presented as quadrant types.
Scientific basis: DISC theory descends from Marston's 1928 book Emotions of Normal People. Marston never developed a formal assessment tool from it; the instruments came later and from other hands, which is why several competing commercial versions now exist with different items, different norms and different item counts. Independent peer-reviewed work is thin: the main published attempt to relate a four-quadrant instrument to a five-factor instrument is a correlational comparison by Jones and Hartley doi:10.19030/ajbe.v6i4.7945, and no meta-analysis of DISC scores against job performance was located for this review. The summary position in the reference literature is that the scientific validity of the DISC assessment has not been demonstrated. Note also that DISC's Conscientiousness label does not mean what Conscientiousness means in the Big Five. See DISC vs Big Five: why four styles aren't enough for a detailed scientific comparison, and Big Five vs DISC vs Belbin for a three-way head-to-head.
Test-retest reliability: Varies by vendor and is mostly reported in vendor technical manuals rather than in independent research.
Predictive validity: Not established for job performance at the level documented for Big Five instruments.
Free to use: Partly. The best-known commercial versions are proprietary and paid, but free DISC-style tests exist: Truity, for example, offers a 38-item DISC assessment with free basic results and an optional paid full report. What is not free anywhere is the underlying methodology, which is not published for independent inspection.
Team-ready: Widely used in training contexts; effective as a communication framework. Not appropriate for high-stakes assessment.
7. Enneagram: mixed evidence, and the nine types are the weakest part
What it measures: Nine personality types, widely used in coaching and spiritual formation contexts.
Scientific basis: Weaker than the Big Five, but the honest verdict is "mixed", not "none". Hook and colleagues systematically reviewed 104 independent samples and found mixed evidence of reliability and validity: some factor analytic work shows partial alignment with Enneagram theory, and subscales relate to other constructs, including the Big Five, in theory-consistent ways. The failures are structural. Factor analyses typically recover fewer than nine factors, no study has used clustering techniques to derive the nine types, and there is little research supporting secondary parts of the theory such as wings and intertype movement doi:10.1002/jclp.23097. In other words, the parts of the Enneagram that behave like continuous traits behave reasonably; the nine-type taxonomy on top of them is the part that has not been demonstrated. For more on myths surrounding personality tests, see five personality science myths that won't die.
Test-retest reliability: Mixed across studies in the systematic review above.
Predictive validity: Not established for job performance.
Free to use: Partially; some versions are free, paid versions exist.
Team-ready: Can generate meaningful personal reflection and discussion, but should be used as a reflective exercise rather than as psychometric data.
Free Personality Tests Compared: Full Ranking Table 2026
| Rank | Tool | Independent peer-reviewed evidence | Open domain | Test-retest reliability | Predictive validity | Free |
|---|---|---|---|---|---|---|
| 1 | IPIP-based assessments | Strong | Yes | ρtt ≈ .82 for Big Five scales | Conscientiousness ρ = .22 for job performance | Yes |
| 2 | Cèrcol | Inherits IPIP evidence | Yes (IPIP) | Inherits IPIP evidence | Inherits IPIP evidence, plus peer ratings | New Moon and First Quarter free, Full Moon paid (free during the open beta) |
| 3 | Open Source Psychometrics | Strong for the IPIP tools | Yes | Inherits IPIP evidence | Inherits IPIP evidence | Yes |
| 4 | 16Personalities | Not located | No | Not located | Not located | Yes |
| 5 | MBTI | Reviews find the evidence insufficient | No | No retest studies in 193-study synthesis | Publisher rules out selection use | No |
| 6 | DISC | Thin | No | Vendor-reported | Not established | Mixed |
| 7 | Enneagram | Mixed across 104 samples | Varies | Mixed | Not established | Partial |
How to Choose the Right Free Personality Test for Your Team
For most teams starting from scratch, a two-step approach works well:
- Begin with a free IPIP-based self-assessment, either via Open Source Psychometrics or First Quarter Cèrcol, to establish a baseline for each person's Big Five profile.
- Add peer assessment (Witnesses) to see where self-perception and others' experience diverge. That gap is often where the most useful development conversations happen, and it is the part with meta-analytic support behind it doi:10.1037/a0021212.
If your team has been using DISC or 16Personalities as a shared language, you do not have to abandon that vocabulary. But layering in a proper Big Five instrument alongside it will give you the dimensional resolution and the published validity evidence those tools lack.
For a deeper dive into what you are actually paying for, or not paying for, when you compare open-source and commercial options, see personality testing: open source vs commercial.
Try a free, sourced team personality assessment
Cèrcol ranks where it does because it is an IPIP implementation, not because it is ours. The items are public domain, the scoring is inspectable, and the numbers on this page are the ones the underlying papers actually report.
First Quarter Cèrcol is 60 items and takes about ten minutes, returning five dimension scores and 30 facet scores. New Moon Cèrcol is a two-minute, ten-item orientation. The Witness peer assessment, part of Full Moon Cèrcol, free to new accounts during the open beta, invites up to twelve people who know you to complete a five-minute forced-choice adjective task. Forced-choice matters for a measurable reason: in a meta-analysis of high-stakes assessment, forced-choice personality measures showed an overall score inflation effect of 0.06, far below the inflation reported for single-statement Likert measures doi:10.1037/apl0000414. See how Cèrcol handles social desirability bias and why 120 items is better than 10 for the measurement design rationale.
Start your team's assessment at cercol.team. Read the scientific foundation to see exactly what you are measuring and why.
Sources
- Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26. https://doi.org/10.1111/j.1744-6570.1991.tb00688.x
- Gnambs, T. (2014). A meta-analysis of dependability coefficients (test-retest reliabilities) for measures of the Big Five. Journal of Research in Personality, 52, 20–28. https://doi.org/10.1016/j.jrp.2014.06.003
- Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78–89. https://doi.org/10.1016/j.jrp.2014.05.003
- Maples-Keller, J. L., Williamson, R. L., Sleep, C. E., Carter, N. T., Campbell, W. K., & Miller, J. D. (2019). Using item response theory to develop a 60-item representation of the NEO PI-R using the International Personality Item Pool: Development of the IPIP-NEO-60. Journal of Personality Assessment, 101(1), 4–15. https://doi.org/10.1080/00223891.2017.1381968
- Gosling, S. D., Rentfrow, P. J., & Swann, W. B., Jr. (2003). A very brief measure of the Big-Five personality domains. Journal of Research in Personality, 37(6), 504–528. https://doi.org/10.1016/S0092-6566(03)00046-1
- McCrae, R. R., & Costa, P. T. (1997). Personality trait structure as a human universal. American Psychologist, 52(5), 509–516. https://doi.org/10.1037/0003-066X.52.5.509
- McCrae, R. R., & Costa, P. T. (1989). Reinterpreting the Myers-Briggs Type Indicator from the perspective of the five-factor model of personality. Journal of Personality, 57(1), 17–40. https://doi.org/10.1111/j.1467-6494.1989.tb00759.x
- Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122. https://doi.org/10.1037/a0021212
- Cao, M., & Drasgow, F. (2019). Does forcing reduce faking? A meta-analytic review of forced-choice personality measures in high-stakes situations. Journal of Applied Psychology, 104(11), 1347–1368. https://doi.org/10.1037/apl0000414
- Erford, B. T., Zhang, X., Sweeting, E., Russo, M., Rashid, A., Sherman, M. F., et al. (2025). A 25-year review and psychometric synthesis of the Myers-Briggs Type Indicator (MBTI), Form M. Journal of Counseling and Development, 103(4), 403–417. https://doi.org/10.1002/jcad.70006
- Pittenger, D. J. (1993). The utility of the Myers-Briggs Type Indicator. Review of Educational Research, 63(4), 467–488. https://doi.org/10.3102/00346543063004467
- Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221. https://doi.org/10.1037/1065-9293.57.3.210
- Hook, J. N., Hall, T. W., Davis, D. E., Van Tongeren, D. R., & Conner, M. (2021). The Enneagram: A systematic review of the literature and directions for future research. Journal of Clinical Psychology, 77(4), 865–883. https://doi.org/10.1002/jclp.23097
- Jones, C. S., & Hartley, N. T. (2013). Comparing correlations between four-quadrant and five-factor personality assessments. American Journal of Business Education, 6(4), 459–470. https://doi.org/10.19030/ajbe.v6i4.7945
- Marston, W. M. (1928). Emotions of Normal People. Kegan Paul, Trench, Trubner and Co.
- IPIP: International Personality Item Pool, public domain statement and item library.
- Open Source Psychometrics Project, raw data archive: https://openpsychometrics.org/_rawdata/
- 16Personalities, NERIS Type Explorer theory page: https://www.16personalities.com/articles/our-theory
- The Myers-Briggs Company, guidance against use in recruitment and selection: https://www.themyersbriggs.com
- Psychometrics Canada, MBTI Step I Form M item count and pricing: https://shop.psychometrics.com
- Truity, free DISC personality test: https://www.truity.com/test/disc-personality-test
- Wikipedia: DISC assessment
- Wikipedia: Myers-Briggs Type Indicator
Further reading
- MBTI vs Big Five: which one should your team use?
- DISC vs Big Five: why four styles aren't enough
- 16Personalities vs Big Five: the viral test that gets it half right
- What reliability and validity mean in personality testing
- Social desirability bias in personality tests
- How personality test scores are calculated: from items to dimensions