Methodology
How scoring works
Every result on iPSYC is built from three decisions: how scores are computed, how we measure reliability, and how we compare you to a reference group. Here's each one in plain language.
1. Scores are per-item means
When you answer a questionnaire, your raw responses are added up within each trait. But different scales have different numbers of items, so a raw total isn't comparable across scales. We divide by the number of items to get a per-item mean — the average answer on the scale's range (usually 1–5).
A score of 3.4 on a 5-point scale means your average response across that trait's items was 3.4. This makes scores directly comparable across scales with different lengths, and it means we can show a single consistent 1–5 scale on every result.
2. Reliability (Cronbach's alpha)
Reliability asks: would you get the same score if you took the test again? We use Cronbach's alpha, a statistic from 0 to 1 that measures how consistently the items in a scale measure the same thing.
- Alpha ≥ 0.7 is generally considered acceptable for research use.
- Alpha ≥ 0.8 indicates good internal consistency.
- Lower alpha doesn't mean the test is broken — short scales (like 2-item BFI-10 facets) naturally have lower alpha because there are fewer items to agree with each other.
We show each scale's alpha on its detail page and results, so you can judge for yourself how much weight to give a score.
3. Norm comparisons
A score of 3.4 means more when you know where it falls in a population. We use percentile ranks: the percentage of people in a reference group who scored at or below your level.
Your result page shows two comparisons where data is available:
- Published norms — from peer-reviewed research on the scale, with the population and source cited.
- Platform aggregates — the running average of all iPSYC users who've taken the same scale, so you can see where you fall among people like you.
Percentiles are computed from the published mean and standard deviation (or platform aggregate) using the standard normal approximation, and marked approximate when the reference sample is small.
Why this matters
Personality measurement is only useful if it's honest. We made three commitments:
- Scores are computed server-side from your raw answers — the client can't change them.
- Every scale is validated before it's shown — item ranges, scoring blocks, and definition structure are checked on import.
- Reliability is shown, not hidden — you can see each scale's alpha and decide how much to trust it.
These are the same standards we use for our research tools.