Skip to main content

Norm Groups in Modern Tests

Learn why classical psychological tests depend on norm groups, why Alva's tests do not, and how results from modern tests should be interpreted.

Written by Malcolm Burenstam Linder

What is a norm group?

A norm group is a reference sample used to give a test score meaning.

In tests based on classical test theory, results are calculated by summing the scores of the individual items and then transforming that sum to a standard scale using the mean and standard deviation of the norm group.

Because the norm group defines the properties of the scale, the quality of the norm group directly determines the quality of the results.


How modern tests work instead

Alva's tests are built on modern test theory, also known as Item Response Theory (IRT).

In IRT, the equivalent of the norm group is the data used to train the statistical model. Rather than scaling a raw sum against a fixed reference sample, the model estimates a set of item parameters for every task in the test:

  • How difficult the task is

  • How well the task differentiates between high and low ability

  • How likely it is that a correct answer was guessed

These item parameters are then used to calculate each candidate's result.

Example: Alva's adaptive logic test

Data from 2,585 individuals was used to estimate the item parameters for 84 tasks in Alva's adaptive logic test. Every result on the logic test is calculated with those parameters taken into account.


Why this matters for result quality

Results from modern tests do not depend on any single norm group. They depend on item parameters estimated from data, which brings two practical advantages:

  • Results improve over time. Item parameters are updated continuously as new data is collected, whereas a norm group is static.

  • The method is current. Item parameters are estimated using machine learning, drawing on current research in statistical modelling, optimisation, and probabilistic programming. Classical test theory rests on statistical theory from the late 19th century.

💡Good to know: Alva uses Bayesian inference in PyMC3 (Salvatier, Wiecki & Fonnesbeck, 2016) to construct scales that match the population of working adults. The models are inspired by those described by Luo and Jiao (2018).


How to interpret Alva's results

Alva's scales are calibrated to the population of working adults. Interpret a candidate's result as a comparison with that population, not with a fixed norm group sample.


References

  • Luo, Y. & Jiao, H. (2018). Using the Stan Program for Bayesian Item Response Theory. Educational and Psychological Measurement, 78(3), 384-408. DOI: 10.1177/0013164417693666

  • Salvatier, J., Wiecki, T. V., & Fonnesbeck, C. (2016). Probabilistic programming in Python using PyMC3. PeerJ Computer Science, 2:55. DOI: 10.7717/peerj-cs.55


Any Questions?

Use the chat in the bottom right corner to connect with a member of our support team.

Did this answer your question?