Back
Likert Scale: Definition, Examples and How to Analyze It
SmartInterview Team

The Short Answer
A Likert item is a single statement with an ordered agreement-style response set, such as "Strongly disagree" to "Strongly agree." A Likert scale, in the original sense, is the sum or average of several such items measuring one underlying construct. The approach is named after Rensis Likert, who introduced summated ratings in the 1930s.
One statement is an item. Several items added together form the scale. Most articles use "Likert scale" to mean a single question, which is why the methodology arguments around it get muddled.
The response set must be balanced and ordered: equal numbers of positive and negative options, with verbal anchors and an optional midpoint.
5 points is the readable default. 7 gives finer discrimination and suits multi-item scales. An even number removes the midpoint and forces a lean, at a cost.
A single item is ordinal. Report the full distribution and top-2-box alongside any average. A summated multi-item score behaves much more like an interval measure in practice.
Reliability statistics like Cronbach's alpha only make sense for a multi-item scale. There is nothing to be internally consistent with when you have one item.
Scale Formats Compared
The format you pick changes what the data can do. Here is the honest trade-off across the five formats you will actually be choosing between.
Format | Discrimination | Midpoint | Mobile readability | Best used for |
|---|---|---|---|---|
4-point forced | Low. Four buckets, no neutral resting place | None. Everyone has to lean | Very good. Short labels, fits one line | Screening, compliance checks, contexts where indifference is not a useful answer |
5-point agreement | Moderate. Enough spread for most tracking | Yes, "Neither agree nor disagree" | Very good. The safest choice on a phone | General-purpose attitude measurement and single standalone items |
7-point agreement | High. Respondents who differentiate get room to | Yes, plus two shades either side of it | Fair. Seven full labels crowd a narrow screen | Multi-item summated scales, psychometrics, sensitive trend detection |
10 or 11-point numeric | Highest on paper, but much of it is noise | Ambiguous. The middle is a number, not a stated position | Good as a row of numbers, poor if you label every point | Likelihood and rating questions, cross-country tracking where wording is hard to translate |
Semantic differential | Moderate to high, depending on point count | Yes, but unlabeled and open to interpretation | Poor. Bipolar adjectives at each end need width | Brand and product image, where the construct is a spectrum rather than agreement |
Note the last row. A semantic differential ("Cheap ... Expensive") is not a Likert item, even though it looks like one in a grid. Likert items are agreement or endorsement ratings of a statement. Mixing the two inside one battery and then summing them is a common and avoidable error.
Labeling: Every Point vs Endpoints Only
Approach | Interpretation | Voice and screen readers | Translation | Space needed |
|---|---|---|---|---|
Fully labeled | More consistent between respondents. Each point means a stated thing | Strong. The options can be read aloud and chosen by name | Hard. Intensity words do not map cleanly across languages | High |
Endpoints only | Reads like a numeric continuum. The middle is left to the respondent | Weak. "Somewhere around four" is awkward to say and to hear | Easier. Only two anchors to localize | Low |
Item vs Scale: The Distinction Most Articles Skip
Rensis Likert's contribution in the 1930s was not the five response options. It was the summated rating: write several statements that all tap the same underlying attitude, have people rate each one on the same ordered response set, then add or average the ratings into a single score for that person.
That total is the Likert scale. A single statement is a Likert item, or a Likert-type item. The difference is not pedantry, because two of the biggest arguments in survey analysis resolve differently depending on which one you have.
Why it changes the ordinal debate
For a single item, the ordinal objection is strong. You have five or seven labeled categories, and nothing guarantees the psychological distance from "Agree" to "Strongly agree" equals the distance from "Neither" to "Agree". Averaging those category numbers means treating unequal steps as equal.
For a summated score across, say, eight items, the objection weakens a lot. The sum takes many more values, its distribution tends toward something reasonably continuous and symmetric, and individual step irregularities partly wash out across items. This is why psychometric practice has long treated multi-item scale scores as suitable for means, correlations, and parametric tests, while being far more cautious about single items.
Why it changes reliability
Internal consistency statistics such as Cronbach's alpha measure how far a set of items hangs together as a measure of one thing. With one item there is no set, so there is nothing to compute. If someone reports an alpha for a single question, something has gone wrong.
Two cautions when you do compute it. First, the common cutoffs you see quoted are conventions, not laws, and they are actively contested in the methodological literature. Second, alpha rises with the number of items, so a long battery of near-duplicate statements can look highly reliable while measuring a very narrow slice of the construct. High alpha is not proof of a good scale.
What a Well-Formed Likert Item Looks Like
A Likert item has two halves, and both have rules.
The statement
One idea only. "The app is fast and easy to use" is double-barreled. A respondent who finds it fast but confusing has no honest answer. Split it.
Declarative, not a question. "Support resolved my issue quickly" works. "How quickly did support resolve your issue?" is a different question type asking for a rating, not agreement.
No negation stacking. "I do not think the checkout is unclear" forces respondents to unpick a double negative before they can answer.
Neutral wording. "The new design is a big improvement" leads. "The new design is an improvement on the previous one" does not.
Concrete and time-bounded where possible. "In the last month, I found it easy to get help" beats "It is easy to get help."
The response set
Balanced. Equal numbers of positive and negative options. "Excellent, Very good, Good, Fair, Poor" is not a balanced scale, it is four shades of positive and one negative, and it will push your results up.
Ordered and monotonic. Each step is clearly more or less than the one before it, in a single direction.
Verbally anchored. At minimum the endpoints. Preferably every point, unless you are running many markets.
Consistent across the battery. Do not switch from a 5-point to a 7-point set halfway through a grid, and do not flip the direction of the scale between screens.
"Don't know" or "Not applicable" sits outside the scale. It is not the midpoint. Place it visually separate and exclude it from the base when you compute means.
Example items
A short battery on a support experience, all rated Strongly disagree / Disagree / Neither agree nor disagree / Agree / Strongly agree:
The agent understood my problem.
I got a clear answer without having to repeat myself.
The time it took to resolve my issue was reasonable.
Reverse-worded: I had to do more work than I should have to get this sorted out.
Those four items measure one construct and can defensibly be summed, provided you reverse-score the last one first. If you are placing items like these in a grid, the layout rules in matrix questions in surveys matter as much as the wording.
How Many Points
There is no universally correct number. There are predictable trade-offs.
5 points
The readable default. Labels fit on a phone, respondents process the set quickly, and the midpoint gives genuinely ambivalent people somewhere honest to go. If you are running one standalone attitude question inside a broader questionnaire, this is usually the right call.
7 points
More discrimination. Respondents who genuinely differentiate get room to express degree, and summated scores built from 7-point items spread out more, which helps when you are trying to detect small movements between waves or between segments. The cost is screen space and cognitive load, especially if you label all seven.
4 points, or any even number
Removing the midpoint forces a lean. Sometimes that is what you want, for instance in a compliance checklist where "neutral" is not a meaningful state. But understand the cost clearly: you do not eliminate ambivalence by removing the button for it. You push people who are genuinely undecided, uninformed, or of two minds onto one side or the other. In agreement scales that pressure tends to land on the agree side, so forced-choice formats can read as more positive than the underlying attitude. If you remove the midpoint, do it because indifference is not a valid answer to your question, not because you dislike seeing a pile in the middle.
Above 7
Expect diminishing returns. Very few people hold attitudes they can reliably sort into nine or eleven distinct verbal grades, and adding points past roughly seven mostly adds noise and layout problems rather than information. Treat that as design guidance rather than a hard finding. The exception is a purely numeric 0-10 rating, which works because respondents read it as a familiar continuum rather than as ten separate labeled positions.
Labeling and the Translation Problem
Fully labeled scales are interpreted more consistently, because every point states what it means instead of leaving respondents to invent a meaning for "4". They are also the only sensible option for voice administration and for screen readers, where a respondent has to hear the options rather than see a row of boxes. If your survey is read aloud, label every point and keep the labels short.
Endpoint-only labeling is more compact, reads as a numeric continuum, and travels better across languages. Its weakness is the middle: two respondents choosing "4" on a 7-point scale may mean quite different things.
Translation is the underrated problem here. Intensity adverbs do not map cleanly between languages. The gap between "agree" and "strongly agree" in English is not guaranteed to equal the gap between the closest available pair in French, German or Japanese, and some languages have no natural single-word equivalent for a mild-agreement anchor. Two consequences follow for multi-country work: use professionally localized anchors rather than machine-translated ones, and be very careful comparing raw means between markets, because part of any gap is the wording rather than the attitude. Endpoint-only or numeric formats reduce, but do not remove, this problem.
Ordinal or Interval? The Honest Answer
Strictly, a single Likert item is ordinal. The categories have a defined order but the spacing between them is not guaranteed to be equal, so the arithmetic operations that a mean assumes are not formally justified.
In practice, treating summated multi-item scores as interval data is standard across applied research, and parametric methods applied to them are generally robust with reasonable sample sizes and non-extreme distributions. The dogmatic position that means are never permissible does not match how the field actually works. The sloppy position, where a 4.17 on a single item gets reported as if it were a measured length, is not defensible either.
The workable middle:
Never publish a mean without the distribution behind it. The mean is a summary of the shape, not a replacement for it. Two very different distributions produce the same average.
Report top-2-box and bottom-2-box alongside it. "62% agree or strongly agree, 11% disagree or strongly disagree" is more actionable than "mean 3.8" and much harder to misread.
For single items, prefer the median and nonparametric tests. Mann-Whitney for two groups, Kruskal-Wallis for more than two, chi-square on the full category distribution when you care about shape rather than central tendency.
For summated multi-item scales, means, t-tests, ANOVA and correlations are reasonable, provided you have checked internal consistency and reverse-scored anything reverse-worded.
Fix your box definitions before you look at the data, and write them into the reporting spec. Top-2-box on a 5-point scale and top-2-box on a 7-point scale are not comparable, so record which one you used.
How to Present Likert Results
Stacked horizontal bars, one bar per item, segments ordered from most negative on the left to most positive on the right, with the counts or percentages labeled. Keep the item order stable between waves. Show the base size on every cut. If you use a diverging layout centered on the midpoint, say so in the chart note, because it changes how the eye reads the balance.
For tracking, plot top-2-box over time rather than the mean. It moves less erratically, it is easier to explain in a business review, and it does not invite the "is 3.6 better than 3.5?" argument that mean-based tracking always produces.
The Rating Tells You Where, Not Why
A Likert battery tells you that agreement with "the pricing is fair" dropped six points this quarter. It does not tell you what changed. That gap is the structural limit of closed-ended rating scales, and it is the reason most serious programs pair them with open questions. See qualitative vs quantitative research for how the two halves fit together.
One practical pattern: trigger a follow-up on the ratings that carry the most information, meaning the low end and the extremes, rather than asking everyone the same generic "why". SmartInterview does this with an AI voice follow-up that fires on a low or extreme rating and probes the reasoning in the respondent's own language, then codes the open answers automatically so the themes are countable next to the scale data. The same multilingual setup helps with the anchor-translation problem above, because the probe adapts to the language the respondent is actually answering in.
Response Biases You Are Measuring Whether You Like It or Not
Acquiescence bias
Some respondents agree with statements regardless of content, particularly when they are tired, uncertain, or trying to be agreeable. Because Likert items are agreement ratings, this bias lands directly on your construct.
The standard remedy is a balanced battery with some items worded in the opposite direction, so that straight-line agreement produces internally contradictory answers you can detect and score out. The remedy has its own cost, and it is a real one. Reverse-worded items are harder to read, respondents misprocess the negation and answer as if it were positive, and the reversed items sometimes cluster together in factor analysis as an artifact of wording rather than of attitude. Use a few reverse-worded items, write them as clean positive statements of the opposite idea rather than as negations, and check afterwards whether they behave like the rest of the set.
Extreme and midpoint response styles
Independent of what they think, some people habitually pick the ends of scales and others habitually stay near the middle. These tendencies vary between individuals and, importantly for international work, systematically between cultures. This is why comparing raw means across countries is risky: part of any difference is scale-use habit, not attitude. Comparing each market against its own history, or standardizing scores within respondent before comparing, is safer than putting country means side by side and declaring a winner.
Social desirability
Statements about behavior people feel judged on, such as effort, honesty, health or spending, attract answers that flatter the respondent. Neutral wording helps, guaranteed anonymity helps more, and self-administered formats generally produce more candid answers than interviewer-administered ones on sensitive topics.
Common Analysis Mistakes
Reporting a single-item mean to two decimals. A 4.17 on five ordered labels implies a precision the instrument does not have. Round it, or report boxes instead.
Comparing means across countries or languages without checking scale use. You may be reporting a translation artifact and a response-style difference as a market insight.
Ignoring a midpoint pile-up. A large middle usually means the statement was vague, irrelevant to that respondent, or asked of people with no basis for an opinion. Fix the item, do not delete the midpoint.
Collapsing categories after seeing the results. Deciding to include "neither" in the positive box because it makes the chart look better is a reporting decision made backwards. Set boxes in advance.
Treating "neutral" as "don't know". They are different states. "Neither agree nor disagree" is a position. "Don't know" is the absence of one, and belongs off the scale and out of the mean base.
Forgetting to reverse-score reverse-worded items. One unreversed item quietly drags the summated score toward the middle and depresses your reliability estimate. Check this before you compute anything.
Summing items that do not measure the same thing. A battery is not automatically a scale. If the items span several constructs, the total is an average of unrelated things and means nothing.
Changing the scale mid-tracker. Moving from 5 points to 7 points, relabeling an anchor, or flipping the direction breaks the trend line. If you must change it, run both versions in parallel for one wave.
Where to Go Next
Lay your items out properly: matrix questions in surveys
Pick the right question type in the first place: single-select questions and multiple choice questions
Write the battery around a satisfaction program: customer satisfaction survey questions
Recover the "why" behind the ratings: how to get 3x more insights using AI surveys
Frequently Asked Questions
What is the difference between a Likert scale and a Likert item?
A Likert item is one statement with an ordered agreement response set. A Likert scale, in the original summated-rating sense, is the total or average of several such items measuring the same underlying construct. Everyday usage blurs the two, but the distinction matters: reliability statistics and parametric analysis are much easier to justify for a summated scale than for a single item.
Should a Likert scale have 5 or 7 points?
Use 5 for standalone attitude questions and for surveys that will mostly be answered on a phone, because five labels stay readable. Use 7 when you are building a multi-item scale, when you need to detect small differences between waves or segments, or when your respondents genuinely differentiate. Above about 7 points you get diminishing returns and more layout problems than extra information.
Should I include a neutral middle option?
Usually yes, if genuine ambivalence is a valid answer to your statement. Removing the midpoint does not remove indifference, it just pushes undecided respondents to one side, and on agreement scales that pressure often inflates agreement. Remove it only when a neutral position is not meaningful for your question. Keep "Don't know" as a separate option outside the scale, never as the midpoint.
Can you calculate a mean from Likert data?
For a summated multi-item scale, yes, and that is standard practice. For a single item it is technically an ordinal measure, so a mean assumes equal spacing you cannot verify. The practical rule is never to report a mean on its own. Publish the full distribution and top-2-box and bottom-2-box next to it, and use the median with nonparametric tests such as Mann-Whitney or Kruskal-Wallis when the analysis rests on a single item.
Is a Likert scale ordinal or interval?
A single item is ordinal, because you cannot assume the psychological gap between "agree" and "strongly agree" equals the gap between "neutral" and "agree". A summated score across several items takes many more values and behaves much more like an interval measure, which is why applied research routinely runs parametric tests on scale scores while being more cautious with individual items.
Why use reverse-worded items?
To catch acquiescence bias, the tendency to agree with statements regardless of content. If some items point the opposite way, a respondent who agrees with everything produces contradictory answers you can detect. The trade-off is that reversed items are harder to read, some respondents miss the reversal, and the reversed items can cluster together as a wording artifact. Use a few, write them as clean opposite statements rather than negations, and always reverse-score them before summing.
How do you analyze Likert data across multiple countries?
Carefully. Verbal anchors do not translate with equal intensity, and response styles differ systematically between cultures, so part of any gap between market means is measurement rather than attitude. Use professionally localized anchors, compare each market against its own trend rather than against other markets, and consider standardizing scores within respondent before making cross-country comparisons.


