TutorTermTell Us What Is StuckGet in Touch
Test selection

There is no list that tells you which test to use, and every list implies there is

Flowcharts get you to a shortlist. Getting from a shortlist to a defensible choice needs your research question, your measurement levels, your design and a look at whether the assumptions actually hold in your data. That last step is the one that gets skipped.

Get your free quote

We reply within 1 business hour, no delays.

No card details required · 100% confidential · On time, or it's free

Four questions, in order

What actually determines the test

First, what are you asking? Whether groups differ, whether variables relate, whether one predicts another, or whether something changed over time. These lead to genuinely different families of test and confusing them is the most fundamental error available.

Second, what are your variables and how are they measured? Nominal, ordinal, interval or ratio, and how many of each. A comparison of two groups on a continuous outcome is a different test from a comparison of two groups on a categorical outcome, and the question can be phrased identically in both cases.

Third, is the design independent or related? Different people in each group, or the same people measured more than once. Getting this wrong is common with repeated-measures designs and it changes the test entirely.

Only then do assumptions come in, and they are what decides between the parametric test and its non-parametric alternative. Checking them is not a formality. It is the justification.

  • Your research question translated into a testable hypothesis
  • Measurement level established for every variable
  • Independent or related design identified correctly
  • Assumptions checked, and an alternative chosen if they fail

What we settle in a session

Usually one session is enough for a specific project

  • The test that answers your question, and why the near alternatives do not
  • Which assumptions matter for that test and how to check each
  • What to do if one fails, with the options ranked
  • How to report the choice in your methodology chapter
  • What an examiner is most likely to challenge about it
The families

Four things you might be asking, and where each leads

Do these groups differ?

Two independent groups points to an independent t-test or Mann-Whitney U. Three or more points to ANOVA or Kruskal-Wallis. The same people measured twice points to a paired t-test or Wilcoxon. Categorical outcomes point to chi-square.

Are these variables related?

Pearson correlation for linear relationships between continuous variables, Spearman for ordinal data or monotonic relationships. Correlation is not causation, and saying so in your limitations is not optional.

Does this predict that?

Linear regression for a continuous outcome, logistic for a binary one, multiple regression for several predictors. This family has the most assumptions and the most ways to go quietly wrong.

Did it change over time?

Repeated-measures ANOVA, mixed models, or their non-parametric equivalents. The recurring mistake is treating repeated measurements from the same participants as if they were independent.

Checking assumptions

What to check, and what to do when it fails

01

Normality, where the test assumes it

Check with plots as well as tests, because significance tests for normality are oversensitive at large samples and underpowered at small ones. If it fails, consider a non-parametric alternative, a transformation, or a robust method.

02

Homogeneity of variance

Levene's test for group comparisons. If it fails, most software offers a correction that does not assume equal variances. Use it and say you did, rather than ignoring the warning.

03

Independence of observations

The assumption you cannot test for and cannot fix afterwards. It is determined by your design: clustered, nested or repeated data needs a model that accounts for the structure.

04

Linearity and multicollinearity, for regression

Scatterplots for the first, variance inflation factors for the second. Highly correlated predictors make individual coefficients unstable and uninterpretable even when the overall model looks strong.

05

Report what you checked

A results chapter that states which assumptions were tested, how, and what followed is much harder to criticize than one where the reader has to assume it happened.

The line we never cross

We help you choose, and we do not help you fish

  • We do not run your analysis for you.
  • We do not rerun tests until one reaches significance.
  • We do not help you drop inconvenient cases to make an assumption hold.
  • We do not write your results or discussion.
  • We teach the reasoning, work through your data with you, and review your write-up.

Removing outliers is sometimes legitimate and sometimes not. The test is whether you decided the rule before you looked at the effect on your result, and whether you report what you removed and why. Read the full policy.

Common questions

Frequently Asked Questions

It depends on whether you are analyzing single items or a summed multi-item scale. Single Likert items are ordinal and usually call for non-parametric tests. A scale summed across several items is commonly treated as continuous, which is widely accepted but is a choice you should state and justify rather than make silently.

Parametric tests are more powerful when their assumptions hold, so they are the default worth trying. Non-parametric alternatives are the fallback when assumptions clearly fail, particularly with small samples or strongly skewed data. Neither is more rigorous in the abstract.

Check how far from normal and with what sample size, because many tests are fairly robust to moderate deviation at larger samples. Then choose deliberately between a non-parametric alternative, a transformation and a robust method, and report the choice. The problem is never non-normality itself, it is non-normality that goes unmentioned.

Yes, but each additional test raises the chance of a false positive, and at some point that needs correcting for, with Bonferroni or something less conservative. Running many tests and reporting only the significant ones is a different matter and is a research integrity problem.

ANOVA, if you have three or more groups. Running every pairwise t-test inflates your error rate substantially. ANOVA tests the overall difference first, and post hoc comparisons then identify where it lies with an appropriate correction.

State your research question, the variables and their measurement levels, the design, the assumptions you checked and how, and the resulting choice, in that order. Two or three clear paragraphs. Most rejected methodology sections are missing the assumptions step rather than the conclusion.

Tell us the question and the variables

One session usually settles test selection for a whole project, including the assumptions to check and the paragraph that justifies it.

Get Help Choosing