Back

Correlation in Survey Data (and Why It Is Not Causation)

Matthieu Saussaye

A correlation coefficient is a single number between -1 and +1 that describes how two variables move together. Positive means they rise together, negative means one rises as the other falls, and zero means no consistent linear pattern. It says nothing about which variable causes the other, or whether either does.



The essentials for survey work:

  • Pearson measures linear association between two continuous variables. It assumes the relationship is a straight line and is sensitive to outliers.

  • Spearman measures monotonic association by correlating ranks instead of raw values. It is the right default for ordinal survey data such as single rating-scale items: Schober, Boer and Schwarte, Anesthesia & Analgesia, 2018 note that it "can be used for ordinal data" and is "relatively robust to outliers".

  • The sign is the direction, the magnitude is the strength. A coefficient of -0.6 is exactly as strong as +0.6.

  • "No correlation" means no linear or monotonic pattern, not "no relationship". A U-shaped relationship can produce a coefficient near zero.

  • Correlation is not causation, and survey data is particularly exposed: confounders, reverse causality and the fact that both answers came from the same person in the same five minutes.

  • Significance depends heavily on sample size. With a large enough sample, a trivially small correlation becomes statistically significant. Report the coefficient itself, not just the p-value.

Which correlation coefficient to use

The choice is driven by the measurement level of your two variables, not by preference.

Coefficient

Data it is for

What it detects

Key assumption or caveat

Pearson r

Two continuous, interval or ratio variables

Linear association

Assumes linearity; distorted by outliers and by restricted range

Spearman rho

Ordinal data, or continuous data that is skewed

Monotonic association, in either direction

Uses ranks, so it discards the size of the gaps between values

Kendall tau

Ordinal data, especially with many tied values

Agreement in ordering between pairs of cases

Handles ties better than Spearman; usually smaller in magnitude

Point-biserial

One continuous variable and one binary variable

Difference in means expressed as a correlation

Mathematically Pearson r applied to a 0/1 variable

Cramer V

Two nominal variables

Strength of association in a cross-tabulation

Runs from 0 to 1 only, so it has no direction

Most survey variables are ordinal, which is why Spearman shows up so often in questionnaire analysis. Schober et al. (2018) set out the same rule: Pearson "is typically used for jointly normally distributed data", while "for nonnormally distributed continuous data, for ordinal data, or for data with relevant outliers, a Spearman rank correlation can be used as a measure of a monotonic association". If your two variables are single items from a rating battery, Spearman is the safer choice than Pearson.

What a correlation coefficient actually measures

Pearson's r summarizes how consistently two variables deviate from their own means in the same direction. If people who score above average on one variable also tend to score above average on the other, r is positive. If above-average on one goes with below-average on the other, r is negative. If there is no consistent pattern, r sits near zero.

Three properties of r are worth holding onto:

  • It is bounded at -1 and +1. Those extremes mean every point falls exactly on a straight line.

  • It is unitless and symmetric. Correlating satisfaction with income gives the same number as correlating income with satisfaction, and changing the units of either variable does not change r.

  • Its square has a direct interpretation. r squared is the proportion of variance in one variable that is shared with the other. An r of 0.30 means about 9% shared variance, which is a useful corrective to how impressive 0.30 sounds.

Pearson versus Spearman, and when ordinal data requires Spearman

Spearman's rho is simply Pearson's r calculated on the ranks of the data rather than the values. That one change alters what the coefficient can detect and what it requires.

Why ordinal survey data pushes you to Spearman

A 5-point scale from "Strongly disagree" to "Strongly agree" is ordinal. The categories have a clear order, but nothing establishes that the distance from "Strongly disagree" to "Disagree" equals the distance from "Agree" to "Strongly agree". Pearson's r takes those distances seriously, because it operates on the numeric codes 1 to 5 as if they were measurements on a ruler. Jamieson, Medical Education, 2004 makes the point about Likert data directly: the categories "have a rank order, but the intervals between values cannot be presumed equal", and the conventional inferential tools for ordinal data are non-parametric ones "such as Chi-square, Spearman's Rho, or the Mann-Whitney U-test".

Spearman only uses the ordering, which is exactly the information an ordinal scale genuinely carries. That is the core argument, and it is why Spearman is the conventional default for correlating two single Likert scale items or two rows from a matrix question.

Practitioners often make an exception for multi-item scales. When several Likert items are summed or averaged into a composite score, the result has many possible values and is commonly treated as interval, so Pearson is widely used. This is a convention with a long-running methodological debate behind it, not a settled fact. Jamieson describes treating ordinal scales as interval as something that "has long been controversial and, it would seem, remains so". On the other side, Norman, Advances in Health Sciences Education, 2010 argues from studies going back to the 1930s that "parametric statistics are robust with respect to violations of these assumptions". The disagreement is real and it is still open. If you use the convention, say so.

The other reasons to prefer Spearman

  • Non-linear but monotonic relationships. If satisfaction rises steeply at first and then flattens, the relationship is monotonic but curved. Pearson underestimates it; Spearman captures it fully.

  • Outliers. One extreme value can drag Pearson's r substantially. Ranking compresses it to just one more position in the order.

  • Skewed distributions. Income, spend and usage counts are usually skewed. Ranks handle that without transformation.

Two caveats. Spearman discards magnitude, so it cannot tell you whether a gap was large or small, only that it existed in a given direction. And when a scale has many tied values, which happens constantly with 5-point items, Spearman needs the tie-corrected formula rather than the simplified shortcut version, and Kendall's tau is often the better-behaved choice.

Reading strength and direction

The sign tells you the direction. Negative correlations are not weaker than positive ones, a point that trips up more reports than it should. A correlation of -0.55 between wait time and satisfaction is a strong finding.

Magnitude is where interpretation gets contextual. Cohen proposed conventional benchmarks for the behavioral sciences. In A Power Primer, Psychological Bulletin, 1992, summarizing the 1988 book, he states that for the significance of a sample r "small, medium, and large ESs are respectively .10, .30, and .50". He is explicit that these are conventions rather than measurements: "the definitions were made subjectively, with some early minor adjustments". They are rules of thumb for a field, deliberately offered as such, and they should be read against your own context. In survey research on human attitudes, correlations above 0.60 between two distinct constructs are uncommon; if you see 0.90, check first whether the two questions are simply asking the same thing in different words.

Whatever the number, do two things before you interpret it. Plot the scatter, because very different data patterns produce identical coefficients. And report a confidence interval, because a point estimate alone hides how much the value could move in another sample. Schober et al. give the same advice: "visual inspection of scatter plots is always advisable, as correlation fails to adequately describe nonlinear or nonmonotonic relationships, and different relationships between variables can result in similar correlation coefficients."

What "no correlation" actually means

A coefficient near zero is a specific, limited statement: there is no consistent linear pattern, for Pearson, or no consistent monotonic pattern, for Spearman, in this sample. It is not a finding of independence.

Four things regularly produce a near-zero coefficient when a real relationship exists:

  • A non-monotonic relationship. If moderate values of one variable go with the highest values of the other, and both extremes go with low values, the up and the down cancel out. The scatter plot shows an obvious inverted U and the coefficient shows nothing.

  • Restricted range. Correlation depends on variation. If you only survey your most loyal customers, satisfaction barely varies among them, and its correlation with anything else is attenuated. Sampling only part of the range can hide a relationship that exists across the full one.

  • Opposing subgroups. A positive relationship in one segment and a negative one in another can average out to zero overall. Correlate within segments before concluding nothing is there.

  • Measurement noise. Unreliable questions add random variation, which pulls observed correlations toward zero. A badly worded item can flatten a genuine association.

The honest way to report a null result is "we found no linear association between X and Y in this sample", followed by the coefficient and its confidence interval. A confidence interval running from -0.05 to +0.35 is not evidence of no relationship; it is evidence that the study could not tell.

Why correlation is not causation, in survey terms

A correlational study observes variables as they naturally occur and measures how they relate. It does not manipulate anything, and manipulation is what licenses causal claims. Schober et al. (2018) state the rule for readers of any correlation: "researchers should avoid inferring causation from correlation". Nearly all survey research is correlational, which means the phrase "drives" should be used with care in every report built from it.

When two survey variables correlate, at least five explanations are live at once.

1. X causes Y

The explanation everyone reaches for first, and one of five.

2. Y causes X, or reverse causality

Support contact frequency correlates with lower satisfaction. Do support interactions make people unhappy, or do unhappy people contact support more? A single cross-sectional survey cannot separate these, because both variables were measured at the same moment.

3. A confounder causes both

This is the big one. Customers who use your mobile app report higher satisfaction. App usage may cause satisfaction, or heavier users may both adopt the app and be more satisfied because they get more value from the product overall. Usage intensity is the confounder, and it inflates the apparent app effect. Adding the confounder as a control variable, through partial correlation or regression, is the standard partial fix. It only works for confounders you thought of and measured.

4. Selection effects

Correlations are computed on the people who answered. If who responds relates to both variables, the correlation is partly a property of your sample rather than your market. Satisfied customers answer satisfaction surveys more readily, which distorts any relationship involving satisfaction. This is one of the reasons your sampling method constrains what a correlation is allowed to mean.

5. Common-method variance

Survey-specific, and widely underestimated. When both variables come from the same person in the same questionnaire, part of their shared variance comes from the person and the instrument rather than from the constructs. Someone in a good mood rates everything higher. Someone who likes your brand rates every attribute favorably, a halo effect. Consistency motives push respondents to answer later questions in line with earlier ones. All of this manufactures correlation between any two attitude items in the same survey.

The practical implication: attitude-to-attitude correlations within a single questionnaire are the weakest evidence in your dataset. Correlations between a survey answer and behavioral data recorded elsewhere are considerably stronger, because the two measurements do not share a source.

What strengthens a causal argument

You cannot get to proof from a cross-sectional survey, but you can get closer.

  • Establish time order. Measure the presumed cause before the presumed effect. A longitudinal or panel design where the same people are surveyed at two points rules out at least the simplest reverse-causality story.

  • Control for what you can measure. Partial correlation and regression remove the influence of confounders you have data for. Be explicit that they cannot remove the ones you do not.

  • Look for a dose-response pattern. If more exposure goes with more effect in a consistent gradient, the causal reading is more plausible than if the relationship appears only at one arbitrary threshold.

  • Use a different data source for one side. Pair survey attitudes with behavioral logs or transaction records to escape common-method variance.

  • Run an experiment when the decision is expensive. A randomized test settles what a correlation only suggests. If a correlation is about to justify a large investment, that is exactly when the experiment is worth it.

  • Ask people directly, then listen. Respondents cannot always report their own causes accurately, but an explanation in their own words often reveals a mechanism, or a confounder, that no correlation matrix would surface. This is where the qualitative side earns its place.

Significance depends on your sample size

A p-value for a correlation tests one narrow hypothesis: that the true correlation in the population is exactly zero. How easily you can reject that hypothesis depends almost entirely on how many respondents you have.

The approximate correlation needed to reach statistical significance at the conventional 5% level, two-tailed, illustrates the point:

Sample size

Roughly the smallest r that reaches p < 0.05

Shared variance at that value

n = 30

About 0.36

About 13%

n = 100

About 0.20

About 4%

n = 400

About 0.10

About 1%

n = 1,000

About 0.06

Under 1%

Two conclusions fall out of that table. In a large sample, a correlation of 0.07 can be reported as statistically significant while explaining well under 1% of the variance, which is almost never a basis for a decision. And in a small sample, a substantively meaningful correlation of 0.30 can fail to reach significance, which is a statement about your sample size rather than evidence that no relationship exists.

Report the coefficient, the confidence interval and the sample size together. "Significant" on its own is close to uninformative.

The correlation matrix trap

Running every variable against every other one is the fastest way to find false positives. Twenty variables produce 190 pairs. At a 5% threshold, roughly ten of them would be expected to come back significant even if nothing in the data were related at all.

If you fish through a matrix, treat what you find as a hypothesis rather than a result. Either adjust for multiple comparisons, or specify which relationships you care about before you look, or hold back part of the sample and check whether the pattern survives in the half you did not explore.

A short checklist before you report a correlation

  1. Check the measurement level. Ordinal items go to Spearman. Composite scales can go to Pearson if you say so.

  2. Plot the scatter. Look for curvature, clusters and outliers before you trust the number.

  3. Report r, n and a confidence interval, not just a p-value.

  4. Square it. Ask whether the shared variance justifies the sentence you were about to write.

  5. List the plausible confounders in the report, including the ones you could not measure.

  6. Choose the verb carefully. "Is associated with" is what the data supports. "Drives" is what an experiment supports.

Correlation is genuinely useful. It narrows the field, points at what to investigate, and turns a large table of closed ended question results into a short list of things worth understanding. What it cannot do is tell you what will happen if you change something.

References

Find the pattern, then ask about it

A correlation matrix tells you two things move together. It never tells you why, and the why is what determines whether acting on it works. Most teams stop at the matrix because going further means running interviews, and interviews are slow.

SmartInterview closes that gap. It runs surveys by voice or text with an AI that probes the reasoning behind an answer, in the respondent's own language, and codes the open responses into themes automatically. You keep the quantitative structure your correlations need, and you get the mechanism behind them from the same study rather than from a separate project three months later.

Put a why behind your next finding: start free or book a conversation.

Frequently Asked Questions

When should I use Spearman instead of Pearson?

Use Spearman when at least one variable is ordinal, such as a single Likert item, when the relationship is monotonic but curved, when there are outliers, or when a distribution is heavily skewed. Spearman correlates the ranks rather than the raw values, so it only relies on ordering. Pearson is appropriate for continuous variables with a roughly linear relationship.

What does the Spearman coefficient of correlation tell you?

It tells you how consistently two variables move in the same direction when both are expressed as ranks. It runs from -1 to +1, where +1 means the two rank orders are identical, -1 means they are exactly reversed, and 0 means no monotonic pattern. Because it uses ranks, it captures direction and consistency but not the size of the gaps between values.

What does "no correlation" mean?

It means the coefficient is near zero, so there is no consistent linear pattern for Pearson, or no consistent monotonic pattern for Spearman, in that sample. It does not mean the variables are unrelated. A curved relationship, a restricted range, opposing subgroups or noisy measurement can all produce a coefficient near zero while a real relationship exists.

What is a correlational study?

A correlational study measures variables as they naturally occur and examines how they relate, without manipulating anything. Almost all survey research is correlational. It can establish that two things go together and how strongly, but because nothing was randomly assigned it cannot rule out reverse causality or an unmeasured third variable causing both.

What counts as a strong correlation in survey data?

It depends on the field and the constructs. Cohen's conventional benchmarks for the behavioral sciences put roughly 0.10 as small, 0.30 as medium and 0.50 as large, and they were offered as rules of thumb rather than fixed rules. In attitude research, values above 0.60 between two genuinely distinct constructs are uncommon. Very high values are often a sign that two questions are measuring the same thing.

Why can't a survey correlation prove causation?

Because nothing was randomly assigned, so at least five explanations remain open: X causes Y, Y causes X, a third variable causes both, the pattern reflects who chose to respond, or both answers share variance simply because the same person gave them in the same questionnaire. Time-ordered measurement, controls, independent data sources and experiments narrow those possibilities; a single cross-sectional survey does not.

Related articles

Sign up for free

Sign up for free

Sign up for free