~/blog

Types of ANOVA

Apr 11, 202610 min readBy Mohammed Vasim
StatisticsMathData Science

You ran a one-way ANOVA comparing three model architectures, got a significant F, and felt good — then a reviewer asks: "Did you consider the interaction between architecture and dataset type? And did you account for the fact that the same folds were used across architectures?" Those two questions correspond to two separate ANOVA designs. Picking the wrong one doesn't just waste power — it can miss the finding entirely.

Taxonomy

Factor: an independent variable that defines groups. In ML: model architecture, dataset type, hyperparameter setting.

Levels: the categories of a factor. Architecture has 3 levels: Transformer, LSTM, CNN.

Between-subjects: different experimental units in each group (one-way ANOVA — different folds per architecture).

Within-subjects: the SAME experimental units contribute to all groups (repeated-measures — same folds evaluated on all architectures).

Fixed effects: factor levels are deliberately chosen. You want conclusions about exactly these levels (e.g., these specific architectures).

Random effects: factor levels are a random sample from a larger population (e.g., random fold splits). You want to generalize beyond the specific levels tested.

ANOVA design taxonomy ANOVA One-Way 1 factor, between Two-Way 2 factors, between Repeated-Measures within-subjects k models, separate folds arch × dataset type same folds, all models Mixed (1 between + 1 within)

One-Way ANOVA

When: one factor, multiple levels, between-subjects design.

ML use case: comparing k architectures evaluated on separate (non-overlapping) folds.

Anchor (one-way)

python
model_A = [0.821, 0.847, 0.835, 0.812, 0.859, 0.828]  # Transformer, x̄=0.834
model_B = [0.791, 0.803, 0.789, 0.812, 0.798, 0.785]  # LSTM, x̄=0.796
model_C = [0.863, 0.879, 0.855, 0.871, 0.868, 0.852]  # CNN, x̄=0.865
# k=3 groups, n=6 per group, N=18 total

ANOVA Table (computed)

Grand mean x̄ = 14.968/18 = 0.8316

SS_between = 6×(0.834−0.832)² + 6×(0.796−0.832)² + 6×(0.865−0.832)² = 6×0.0000044 + 6×0.0012406 + 6×0.0010963 = 0.01405

SS_within = SS_A + SS_B + SS_C = 0.001483 + 0.000503 + 0.000513 = 0.002499

SourcedfSSMSF
Between (arch)20.014050.00702442.1
Within (error)150.0024990.000167
Total170.016549

F(2,15) = 42.1, p < 0.0001. Reject H₀. At least one architecture differs.

Tukey HSD Post-Hoc

HSD = q(0.05, k=3, df=15) × √(MS_within/n) = 3.67 × √(0.000167/6) = 3.67 × 0.00527 = 0.01934

| Pair | |x̄ᵢ − x̄ⱼ| | HSD | Significant? | |------|----------|-----|-------------| | Transformer vs LSTM | 0.0373 | 0.0193 | Yes | | Transformer vs CNN | 0.0310 | 0.0193 | Yes | | LSTM vs CNN | 0.0683 | 0.0193 | Yes |

All three architectures differ significantly.

Two-Way ANOVA

When: two factors simultaneously, between-subjects. Each factor-level combination is a "cell."

ML use case: accuracy depends on architecture AND dataset type, and you want to know if the architecture effect is consistent across dataset types (interaction).

Factor A: Architecture (Transformer, LSTM, CNN) — 3 levels. Factor B: Dataset type (clean, noisy) — 2 levels.

Cell Means Table

Clean dataNoisy dataRow mean
Transformer0.8710.8230.847
LSTM0.8090.7720.790
CNN0.8780.8410.860
Column mean0.8530.8120.832

Three Tests in One

1. Main effect of Factor A (architecture): does mean accuracy differ across architectures, averaged over dataset type?

2. Main effect of Factor B (dataset): is clean data significantly better than noisy, averaged over architectures?

3. Interaction effect A×B: does the architecture ranking change depending on dataset type? If CNN is best for clean but not for noisy — that is an interaction.

SS Decomposition

SS_total = SS_A + SS_B + SS_AB + SS_within

SourceSSdf
Factor A (arch)SS_Aa−1 = 2
Factor B (data)SS_Bb−1 = 1
Interaction A×BSS_AB(a−1)(b−1) = 2
Within (error)SS_Wab(n−1) = 6

Interaction Interpretation

If interaction is significant: the architecture effect changes depending on dataset. You cannot make a single architecture recommendation without specifying the dataset. CNN may be best for clean data but not noisy data.

If interaction is NOT significant: the factors act independently. You can recommend the best architecture regardless of dataset.

Interaction plot: parallel lines = no interaction; crossing = interaction 0.770 0.800 0.840 0.880 Clean Noisy Transformer LSTM CNN Parallel lines → no interaction. All architectures suffer equally on noisy data.

Parallel lines here indicate no significant interaction — each architecture's performance drops by roughly the same amount from clean to noisy data. The architecture ranking is stable: CNN > Transformer > LSTM regardless of dataset type.

Repeated-Measures ANOVA

When: the same experimental units (folds) are measured under all conditions.

ML use case: the same 4 CV folds are evaluated on all 3 architectures. Because the same fold data is used across architectures, observations within a fold are correlated. Ignoring this correlation uses one-way ANOVA and inflates residual variance.

Why It Is More Powerful

In one-way ANOVA, fold-to-fold variability (some folds are just harder) ends up in SS_within, inflating the denominator. Repeated-measures ANOVA partitions out that fold-level variance:

SS_total = SS_between_subjects + SS_within_subjects SS_within_subjects = SS_treatment + SS_error

Removing SS_between_subjects (fold-to-fold differences) from the error term gives a smaller denominator → larger F → higher power.

Between-subjects vs within-subjects (repeated-measures) design Between-subjects (one-way) Folds 1-4 Folds 5-8 Folds 9-12 Transformer LSTM CNN Each model sees different folds Fold variance → error term Within-subjects (repeated-measures) Same Folds 1-4 Trans LSTM CNN Each fold scored on all models Fold variance partitioned out → higher power

Sphericity Assumption

Repeated-measures ANOVA requires sphericity: the variances of differences between all pairs of conditions must be equal.

  • Mauchly's test: H₀ = sphericity holds. If p < 0.05, apply correction.
  • Greenhouse-Geisser (GG): conservative. Use when ε < 0.75.
  • Huynh-Feldt (HF): less conservative. Use when ε ≥ 0.75.

Both corrections reduce the effective df, making the test more conservative. ε = 1.0 means sphericity holds perfectly; ε < 1.0 means violations, and df is multiplied by ε.

Decision Framework

DesignFactorsObservationsML Use Case
One-Way1Independent (between)Compare k architectures on separate folds
Two-Way2Independent (between)Architecture × dataset size; arch × regularization
Repeated-Measures1+Correlated (within)Same folds evaluated on all models
Mixed ANOVA1 between + 1 withinBothArchitecture (between) × training stage (within)

Key question: are the same experimental units (folds, subjects, users) measured under multiple conditions? If yes → repeated-measures or mixed.

Post-Hoc Tests

Every significant ANOVA must be followed by a post-hoc test. ANOVA only tells you "at least one group differs."

TestControlsConservative?Use When
Tukey's HSDFWER (all pairs)ModerateAll pairwise comparisons
BonferroniFWER (specific pairs)MostSpecific comparisons planned in advance
SchefféFWER (all contrasts)Most conservativeCustom contrasts (e.g., A vs average of B+C)

Use Tukey's HSD by default. Bonferroni is more conservative than Tukey for all-pairs comparisons — use it only when testing specific pre-planned comparisons. Scheffé handles arbitrary contrasts, not just pairwise.

Code

python
import numpy as np
from scipy import stats
from statsmodels.stats.multicomp import pairwise_tukeyhsd
import pandas as pd

# One-way ANOVA anchor
model_A = [0.821, 0.847, 0.835, 0.812, 0.859, 0.828]
model_B = [0.791, 0.803, 0.789, 0.812, 0.798, 0.785]
model_C = [0.863, 0.879, 0.855, 0.871, 0.868, 0.852]

f_stat, p_val = stats.f_oneway(model_A, model_B, model_C)
print(f"One-way ANOVA: F={f_stat:.3f}, p={p_val:.6f}")

# Manual SS decomposition
all_data = model_A + model_B + model_C
grand_mean = np.mean(all_data)
groups = [model_A, model_B, model_C]
SS_between = sum(len(g) * (np.mean(g) - grand_mean)**2 for g in groups)
SS_within = sum(sum((x - np.mean(g))**2 for x in g) for g in groups)
df_between, df_within = len(groups) - 1, len(all_data) - len(groups)
MS_between = SS_between / df_between
MS_within = SS_within / df_within
print(f"SS_between={SS_between:.6f}, df={df_between}, MS={MS_between:.6f}")
print(f"SS_within={SS_within:.6f}, df={df_within}, MS={MS_within:.6f}")

# Tukey's HSD post-hoc
scores = model_A + model_B + model_C
labels = ['A']*6 + ['B']*6 + ['C']*6
tukey = pairwise_tukeyhsd(scores, labels, alpha=0.05)
print("\nTukey's HSD post-hoc:")
print(tukey)
text
One-way ANOVA: F=42.085, p=0.000001

SS_between=0.014048, df=2, MS=0.007024
SS_within=0.002499, df=15, MS=0.000167

Tukey's HSD post-hoc:
 Multiple Comparison of Means - Tukey HSD, FWER=0.05
 ===================================================
 group1 group2 meandiff p-adj   lower   upper  reject
 -------------------------------------------------------
      A      B  -0.0373  0.0001 -0.0499 -0.0248  True
      A      C   0.0310  0.0001  0.0184  0.0436  True
      B      C   0.0683  0.0001  0.0558  0.0809  True
 -------------------------------------------------------

Test Your Understanding

  1. In a two-way ANOVA on architecture × dataset type, the interaction p-value is 0.032. You observe that CNN is best for clean data but third-best for noisy data. Can you report "CNN is the best architecture"? Why or why not?

  2. One-way ANOVA gives F=42 for these architectures, but repeated-measures ANOVA (using the same folds for all models) gives F=95. Why is the repeated-measures F larger for the same data? What variance is being removed in each case?

  3. Sphericity is violated (Mauchly's p=0.02) with ε=0.60. You choose between Greenhouse-Geisser and Huynh-Feldt corrections. Which should you use, and what does it do to the degrees of freedom in your F-test?

  4. Your ANOVA F is significant (p=0.003) with k=5 architectures. You want to test whether Transformer is significantly different from the average of all LSTM variants combined. Which post-hoc test handles this contrast, and why can't Tukey's HSD do it?

  5. You design an experiment with 3 architectures (between) × 4 training epochs (within). How many cells does this mixed ANOVA have? What is the key interaction effect, and what would it mean practically if that interaction is significant?

Comments (0)

No comments yet. Be the first to comment!

Leave a comment