Home Statistical Dictionary Effect Size
Effect Size
A standardized measure of the magnitude of a difference or relationship, independent of sample size, that complements a hypothesis test's p-value.
In Plain English
A p-value tells you whether an effect is statistically detectable, but it says nothing about how big that effect actually is, a tiny, practically meaningless difference can produce a significant p-value with a large enough sample. Effect size fills that gap: it's a standardized number that describes how large or important a difference or relationship is, in units that let you compare across different studies, scales, and even different types of statistical tests.
Definition
Effect size is any of a family of standardized measures quantifying the magnitude of a phenomenon, a difference between group means, the strength of a relationship, or the size of an association, independent of sample size, in contrast to a p-value, which conflates effect magnitude with sample size and is not itself a measure of practical importance. Different families of effect size measures suit different data structures: standardized mean differences (Cohen's d, Hedges' g, Glass's delta) for comparing two group means, variance-explained measures (eta squared, partial eta squared, omega squared) for ANOVA designs, and association measures (phi coefficient, Cramer's V) for categorical data. Reporting effect size alongside a p-value is now standard practice in most fields, since it allows readers to judge practical significance and enables meta-analyses to combine results across studies measured on different scales.
Properties
- Effect size and statistical significance are logically independent, a very small, practically unimportant effect can be statistically significant with a large enough sample, while a genuinely large effect can fail to reach significance with a small sample, this is the central reason effect size reporting has become standard alongside p-values.
- Effect size measures are the fundamental input to meta-analysis, which combines results across multiple studies that may have used different sample sizes, measurement scales, or even different specific tests, a standardized effect size (like Cohen's d) puts these disparate results on a common, combinable scale in a way raw p-values or unstandardized differences cannot.
- Conventional benchmarks for 'small,' 'medium,' and 'large' effect sizes, originally proposed by Jacob Cohen for measures like Cohen's d and eta squared, are widely used but are explicitly rules of thumb rather than universal truths, what counts as a practically important effect size genuinely depends on the specific field and research question.
At a Glance
Related Terms