What is a norm group?
A norm group is a reference sample used to give a test score meaning.
In tests based on classical test theory, results are calculated by summing the scores of the individual items and then transforming that sum to a standard scale using the mean and standard deviation of the norm group.
Because the norm group defines the properties of the scale, the quality of the norm group directly determines the quality of the results.
How modern tests work instead
Alva's tests are built on modern test theory, also known as Item Response Theory (IRT).
In IRT, the equivalent of the norm group is the data used to train the statistical model. Rather than scaling a raw sum against a fixed reference sample, the model estimates a set of item parameters for every task in the test:
How difficult the task is
How well the task differentiates between high and low ability
How likely it is that a correct answer was guessed
These item parameters are then used to calculate each candidate's result.
Example: Alva's adaptive logic test
Data from 2,585 individuals was used to estimate the item parameters for 84 tasks in Alva's adaptive logic test. Every result on the logic test is calculated with those parameters taken into account.
Why this matters for result quality
Results from modern tests do not depend on any single norm group. They depend on item parameters estimated from data, which brings two practical advantages:
Results improve over time. Item parameters are updated continuously as new data is collected, whereas a norm group is static.
The method is current. Item parameters are estimated using machine learning, drawing on current research in statistical modelling, optimisation, and probabilistic programming. Classical test theory rests on statistical theory from the late 19th century.
đĄGood to know: Alva uses Bayesian inference in PyMC3 (Salvatier, Wiecki & Fonnesbeck, 2016) to construct scales that match the population of working adults. The models are inspired by those described by Luo and Jiao (2018).
How to interpret Alva's results
Alva's scales are calibrated to the population of working adults. Interpret a candidate's result as a comparison with that population, not with a fixed norm group sample.
References
Luo, Y. & Jiao, H. (2018). Using the Stan Program for Bayesian Item Response Theory. Educational and Psychological Measurement, 78(3), 384-408. DOI: 10.1177/0013164417693666
Salvatier, J., Wiecki, T. V., & Fonnesbeck, C. (2016). Probabilistic programming in Python using PyMC3. PeerJ Computer Science, 2:55. DOI: 10.7717/peerj-cs.55
Any Questions?
Use the chat in the bottom right corner to connect with a member of our support team.
