Open the methods chapter of almost any questionnaire-based paper and you'll see the same sequence of analyses, in the same order. It's not a coincidence — structural equation modeling has a required build-up. You can't jump straight to testing your hypotheses; you first have to prove your survey measured what it claims to.
SEM is really two models stacked together, as the path diagram above shows: a measurement model that links your survey items to the latent constructs behind them, and a structural model that tests the relationships among those constructs. The pipeline exists to validate the first before trusting the second.
Note
The Pipeline at a Glance
| Step | Purpose | What a paper reports |
|---|---|---|
| 1. Descriptive stats | Sample and item profile | Means, SD, skew/kurtosis |
| 2. Correlation | Inter-construct relationships | Correlation matrix |
| 3. Reliability | Internal consistency | Cronbach's α, composite reliability |
| 4. EFA | Discover the factor structure | Loadings, variance explained |
| 5. CFA | Confirm the measurement model | AVE, discriminant validity, fit indices |
| 6. SEM | Test the hypotheses | Path coefficients, model fit |
| 7. Mediation / moderation | Mechanisms and conditions | Indirect effects, interactions |
Each step is a gate. If reliability fails, you fix the scale before EFA. If CFA fit is poor, you don't report the SEM. The order protects you from building conclusions on a shaky measure.
Steps 1–3: Describe, Correlate, Check Reliability
The groundwork. Descriptive statistics profile the sample and check that items aren't badly skewed. Correlations show how the constructs relate and flag multicollinearity. Reliability confirms the items in each scale hang together.
- Descriptive: means, standard deviations, and skew/kurtosis for normality
- Correlation: the inter-construct matrix, and a check for redundancy
- Reliability: Cronbach's α (≥ 0.7 acceptable) and composite reliability (CR ≥ 0.7)
Tip
Step 4: EFA — Discover the Structure
Exploratory factor analysis asks how many latent factors the items actually reflect, and which items load on which factor. It's where a messy battery of questions resolves into a small number of clean constructs — or reveals that an item belongs somewhere you didn't expect.
- Number of factors — often chosen by eigenvalues or a scree plot
- Factor loadings — how strongly each item ties to its factor
- Cross-loadings — items that load on more than one factor are candidates to drop
Step 5: CFA — Confirm the Measurement Model
Confirmatory factor analysis takes the structure EFA suggested and tests whether it actually fits the data. This is where validity is established and where most papers live or die, because CFA reports the fit indices reviewers scrutinize.
| Criterion | Common threshold |
|---|---|
| Convergent validity (AVE) | AVE ≥ 0.5 |
| Composite reliability (CR) | CR ≥ 0.7 |
| Discriminant validity | √AVE greater than inter-construct correlations |
| Model fit — CFI / TLI | ≥ 0.90 (≥ 0.95 ideal) |
| Model fit — RMSEA | ≤ 0.08 (≤ 0.06 ideal) |
| Model fit — SRMR | ≤ 0.08 (≤ 0.05 ideal) |
Watch out
Step 6: SEM — Test the Hypotheses
With a validated measurement model, the structural model estimates the paths between constructs — your actual hypotheses. Each path is the β on the arrow in the diagram above: how much one latent construct drives another, with a significance test and its own model fit.
- Path coefficients (standardized β) and their significance for each hypothesis
- Overall structural model fit, reported the same way as CFA
- R² for each endogenous construct — how much the model explains
Step 7: Mediation and Moderation
Most interesting hypotheses aren't a single arrow. Mediation asks whether A affects B through a third construct; moderation asks whether the A→B path changes under a condition. These extend the structural model into the mechanisms a study usually cares about most.
Tip
The SEM Pipeline in the SKARI Statistical Lab
SKARI's Statistical Lab runs the whole questionnaire pipeline — the same sequence, with the exact statistics a paper reports, in one place.
- Descriptive statistics and correlation for the groundwork
- Reliability with Cronbach's α and composite reliability
- Exploratory and confirmatory factor analysis (EFA, CFA)
- CFA validity and fit indices — AVE, discriminant validity, CFI, TLI, RMSEA, SRMR
- Structural equation modeling with standardized paths and model fit
- Mediation and moderation with bootstrap confidence intervals
Takeaway
Frequently Asked Questions
Do I need both EFA and CFA?
For a new or adapted scale, yes — EFA to discover the structure, CFA to confirm it (ideally on a separate sample). For a well-established scale, CFA alone is often enough.
What fit indices should I report?
Commonly CFI and TLI (≥ 0.90), RMSEA (≤ 0.08), and SRMR (≤ 0.08), alongside the model chi-square. Report several, not just the one that looks best.
Can I skip straight to SEM?
No — without validating the measurement model first, a good-looking structural result may rest on constructs your survey never measured cleanly.
How large a sample do I need?
SEM is data-hungry; rules of thumb range from 10 cases per estimated parameter to a few hundred minimum, depending on model complexity.
Key Takeaways
It's a
Pipeline
not one test
First
Measure
EFA + CFA
Then
Structure
the SEM paths
Report
Fit
CFI/RMSEA/SRMR
Survey-based SEM is a disciplined sequence: describe, correlate, confirm reliability, discover structure with EFA, validate it with CFA, then test hypotheses with the structural model and its extensions. Validate the measurement before you trust the paths — that order is what makes the conclusions publishable.
Takeaway
Structural Analysis
The methods behind the pipeline
Descriptive Statistics
Where the pipeline starts
Correlation ≠ Causation
Reading construct relationships