How can we tell whether two categorical variables move together or merely look different by chance?
An online store wants to know whether checkout completion differs by payment method. The outcomes of 600 customers are cross-tabulated below.
Observed and expected counts
If payment method and checkout completion were unrelated, how many customers would we expect in each cell?
| Completed | Abandoned | Total | |
|---|---|---|---|
| Card | 180 | 120 | 300 |
| Digital wallet | 240 | 60 | 300 |
| Total | 420 | 180 | 600 |
Completion is 80% with a digital wallet and 60% with a card. The question is whether that pattern is stronger than we would expect from sampling variation alone.
Key question
How far are the observed counts from the counts expected if the two variables were independent?
How do observed and expected counts differ?
The null hypothesis for a chi-square test of independence says that the two categorical variables are unrelated. Expected counts are calculated from the row totals and column totals.
Expected count
(row total × column total) ÷ grand total
For card users who completed checkout, the expected count is 300 × 420 ÷ 600 = 210. The observed count is 180.
How is the χ² statistic built?
For each cell, the observed-minus-expected difference is squared and divided by the expected count. The cell contributions are then added.
How each cell contributes to χ²
Larger gaps between observed and expected counts contribute more.
| Cell | Observed O | Expected E | (O−E)²/E |
|---|---|---|---|
| Card × completed | 180 | 210 | 4.29 |
| Card × abandoned | 120 | 90 | 10.00 |
| Wallet × completed | 240 | 210 | 4.29 |
| Wallet × abandoned | 60 | 90 | 10.00 |
Reading χ²
A larger χ² means the observed table is farther from the pattern expected under independence. The p-value, calculated with the degrees of freedom, determines statistical significance.
What does a significant result tell us?
Chi-square test of independence
Payment method × checkout completion
χ²
28.57
df
1
p-value
< .001
Cramér’s V
0.22
A small p-value leads us to reject independence. It does not establish that payment method caused the difference in checkout completion.
Association is not causation
Customers who choose digital wallets may already differ in purchase intent. The test detects association, not a causal mechanism.
With a very large sample, even a small difference in proportions can be significant. Cramér’s V is commonly used to describe the strength of association.
When should the test be used?
| Test | Question | Example |
|---|---|---|
| Goodness of fit | Does one categorical variable follow a stated distribution? | Are four brands preferred equally? |
| Independence | Are two categorical variables associated? | Is payment method related to completion? |
A non-significant result
p = .31 does not prove independence. It means the current table does not provide enough evidence against it.
Change the cell counts
Adjust the four cells below. The expected counts, χ² statistic, p-value, and Cramér’s V update together.
χ²
28.57
p
<0.0001
Cramér's V
0.218
Total n
600
Expected counts
210.0 · 90.0 · 210.0 · 90.0
Key lesson
The chi-square test evaluates how different the pattern of category proportions is from the expected pattern.
The chi-square test asks whether the pattern of proportions differs, not merely whether the raw counts differ.