Back

Rank Order Questions: How to Use Them Without Losing Respondents

SmartInterview Team

A rank order question asks respondents to put a list of items in order of preference, importance or priority. Unlike a rating scale, it forces a trade-off: respondents cannot call everything important. That makes ranking the right tool when you need to know what comes first, and the wrong tool when you need to know how much more one item matters than another.



The short version of what you need to decide before you use one:

  • Use ranking when priority is the question. "Which of these features should we build first?" is a ranking question. "How satisfied are you with support?" is not.

  • Keep full rankings short. Roughly 5 to 7 items is a sensible working ceiling. Past that, respondents order the top and bottom and guess at the middle.

  • Prefer "rank your top 3" over ranking a long list. Partial ranking collects the part of the answer respondents actually know, and drops the noise.

  • Watch the format. Drag-and-drop looks elegant on desktop and breaks down on touch screens, screen readers and keyboard navigation. Numbered entry is uglier and more robust.

  • For 15 or more items, use MaxDiff instead. Best-worst scaling scales where ranking does not, and produces scores you can actually compare.

  • Ranking data is relative, not absolute. You learn order. You do not learn distance, and you do not learn whether the top-ranked item matters to that person at all.

Ranking vs the alternatives

Ranking is one of several ways to measure priority. Each one buys you something different, and the cost is usually paid in respondent effort.

Method

What you learn

Respondent effort

Practical item ceiling

Best used for

Full ranking

Complete order of every item, no distances

High, and rises steeply with each added item

About 5 to 7

Short lists where every position matters

Top-N ranking

Order of the leading few, plus who never gets picked

Moderate

Around 8 to 12 shown, 3 ranked

Longer lists where only the leaders matter

Rating battery

Absolute level per item, comparable across people

Low per item

High, but attention fades

Tracking levels over time

MaxDiff

Utility scores on one scale, ratio-like comparisons

Low per screen, many screens

20 to 30 or more

Prioritizing long lists of features, messages or claims

Constant sum

Order and relative magnitude, as an allocation

High, arithmetic required

About 4 to 6

Budget or attention splits where magnitude matters

The row that trips teams up most is the rating battery. It is the cheapest to field and the easiest to analyze, which is exactly why it gets used for questions it cannot answer.

What ranking measures that rating does not

Ask ten people to rate eight product features on a 5-point importance scale and you will usually get a wall of 4s and 5s. Nobody wants to say a feature is unimportant. Security is important. Speed is important. Price is important. The battery comes back with six items averaging between 4.2 and 4.5, and you have learned nothing you can act on.

This is scale-use bias, and the resulting lack of discrimination is the single most common failure of importance research. Respondents are not lying. Each item genuinely is important in isolation. The question simply never asked them to choose.

Ranking removes that escape route. A respondent can only put one item first. Putting security first means putting price second, and that trade-off is the information you wanted. It also cancels out a lot of individual scale habits: the person who rates everything 5 and the person who rates everything 3 can still produce the same ranking, so ranking is naturally immune to the response styles that distort Likert scale batteries and long matrix questions.

The cost: ranking data is ipsative

Everything ranking gains from forcing a choice, it pays for by producing relative data. Ranks are ipsative, meaning each respondent's answers are defined only in relation to their own other answers.

Three consequences follow, and all three are routinely ignored in reporting:

  • You get order, not distance. Items ranked 1 and 2 may be nearly tied, or the first may be overwhelmingly more important. The rank numbers look identical either way.

  • You get no absolute level. Someone's top-ranked item might still be something they barely care about. Ranking always produces a winner, even in a list where nothing matters to that person.

  • Ranks are not freely comparable across respondents. A rating of 4 means roughly the same thing for two people. A rank of 2 means "second out of this person's own set", which is a different statement each time.

The practical fix is to pair the methods rather than choose between them. Run a short rating battery for absolute levels and a ranking question for priority. If the rating battery says everything is important and the ranking says one item is consistently first, you know which one to fund.

Formats, and which one to actually use

Drag-and-drop

The default in most modern survey tools. On a desktop with a mouse it is genuinely the most intuitive interface: respondents see the whole list, move items, and understand the result immediately.

The problems appear everywhere else. On a phone, dragging an item competes with scrolling the page, so respondents scroll when they meant to drag and drag when they meant to scroll. Long lists that do not fit on one screen make it worse, because the list has to auto-scroll while an item is held. For screen reader users and anyone navigating by keyboard, drag-and-drop is often either unusable or supported by a hidden alternative interface that is not well tested. Given that most survey traffic now arrives on mobile, treating drag-and-drop as the safe default is backwards.

Numbered entry

Respondents type or select a rank number next to each item. It looks dated and it requires validation logic to catch duplicate numbers and gaps. It is also robust: it works with a keyboard, works with assistive technology, works on any screen size, and does not silently lose an answer because a drag ended in the wrong place. For long lists and mobile-heavy samples, it is the more reliable choice.

Pick your top N

Respondents select their top 3 in order by tapping items in sequence. This is the format that fails least often, mostly because it asks for less. It maps cleanly to a tap interface and produces exactly the part of the ranking you can trust.

Paired comparison

Show two items, ask which one wins, repeat. Each judgment is trivially easy and the data is high quality. The problem is volume: the number of possible pairs grows quadratically, so a full paired design becomes impractical well before you reach a dozen items. MaxDiff exists largely to solve this.

The cognitive load ceiling

Ranking effort does not grow in line with the number of items. It grows much faster, because ordering a list means holding items in memory and comparing them against each other, not evaluating each one on its own.

Respondents also do not rank uniformly well across the list. They usually know their favorite and their least favorite with real confidence. The extremes are stable. The middle is not, and if you re-ask the same person the same ranking a week later, the middle positions are the ones most likely to move. For a long list, positions 5 through 10 are close to noise dressed up as data.

A practical working guide, not a research threshold:

  • Up to about 5 to 7 items: a full ranking is reasonable. Treat this as a rule of thumb rather than a validated cutoff.

  • 8 to 12 items: ask for the top 3, and consider also asking for the single worst item.

  • More than about 15 items: stop ranking. Use MaxDiff.

Long, effortful questions are also one of the reasons survey response rates keep falling. A mandatory 12-item drag-and-drop ranking placed early in a questionnaire is a good way to lose the respondents you paid for.

Partial ranking as the pragmatic default

Top-N ranking asks respondents to order only their leading few items and leave the rest untouched. It is faster, it is easier on mobile, and it collects the portion of the answer respondents can actually give reliably.

It does change your analysis. Unranked items are not tied at the bottom, and they are not missing data either. The correct reading is "this respondent did not place this item in their top 3", which is a real and useful signal. Report the percentage of respondents who ranked each item first and the percentage who placed it in their top 3, and let items nobody picks show up as exactly that. Do not impute an average rank for unranked items. That invents a number the respondent never gave you.

MaxDiff: the scalable alternative

MaxDiff, also called best-worst scaling, is what you use when the list is too long to rank. Instead of one big ordering task, respondents see a series of small sets, typically four or five items at a time, and answer two easy questions per screen: which is best, and which is worst.

The sets are drawn from an experimental design so that every item appears a balanced number of times and alongside a balanced mix of other items. Across all screens, each respondent has effectively made a large number of comparisons without ever facing a long list.

Why it scales

Each screen is short regardless of how many items are in the study. Adding items adds screens, not difficulty. That is why MaxDiff comfortably handles 20 to 30 items or more, where a full ranking collapses well before 15.

Picking the best and the worst also extracts more information than picking a favorite alone. It pins down both ends of the set in one pass.

What it produces

The choices are fed into a scoring model, usually a form of choice model estimated per respondent, which returns utility scores placed on a common scale. Rescaled so the scores sum to 100, those numbers support the comparisons ranking cannot make: you can say one item scores roughly twice another, and you can compare items across segments because everyone sits on the same scale.

What it costs

  • More screens. Covering the item list properly takes a series of sets, not one question. Respondents find each screen easy, but the block is longer.

  • A real design. You need a balanced experimental design, not randomly assembled sets. Getting this wrong biases the scores.

  • A scoring model. Raw best-minus-worst counts are a usable first pass, but proper utilities require estimation, which means either a platform that does it for you or an analyst who does.

  • Sample requirements. Because each respondent sees only a subset of comparisons, you need enough respondents to cover the design.

MaxDiff is heavier machinery. Do not reach for it to prioritize five items. Reach for it when you have a genuinely long list of features, messages or claims and the whole point of the study is to find the top handful.

How to analyze ranking data

Ranking data is ordinal. That single fact should shape every number you report.

Start with the distribution

Before any summary statistic, look at how often each item lands in each rank position. This is the most honest view of the data and it exposes patterns a single average hides. An item ranked first by a large minority and last by everyone else is a polarizing item, and it will look mediocre and forgettable if you only report its mean.

Percent ranked first, and percent in the top 3

These two are usually the most decision-ready numbers you have. They are easy to explain to stakeholders, they map directly onto the choice you are making, and they do not pretend the ranks are interval measurements. Report both, because they answer different questions: percent ranked first identifies the leader, percent in the top 3 identifies broad acceptability.

Mean rank, handled carefully

Mean rank is common and it is not useless, but it assumes the gap between ranks 1 and 2 equals the gap between ranks 4 and 5. Ranks carry no such guarantee. Use mean rank as a rough ordering device, always show it alongside the distribution, and never present it as a measure of how much more important one item is than another. Median rank is often the safer summary. With partial rankings, mean rank is worse still, because it can only be computed on the respondents who ranked that item, which quietly changes the base for every item.

Rank-sum and points-weighted scores

Many tools report a weighted score: first place gets 5 points, second gets 4, and so on. It produces a clean league table and stakeholders like it. Be aware that the weights are arbitrary. Nothing in the data says first place is worth exactly one more point than second. Switch to a steeper weighting and the ordering near the top can change. If you use a points score, state the weights in the report and check whether your conclusion survives a different scheme.

Appropriate tests

Use nonparametric methods. The Friedman test handles differences across ranked items within the same respondents. Wilcoxon signed-rank works for comparing two items. Mann-Whitney or Kruskal-Wallis compare an item's ranks between independent groups such as segments. Spearman's correlation compares two rank orderings. Running a t-test or ANOVA directly on raw ranks treats ordinal positions as interval scores, which is the same mistake as reporting mean rank without caveats, just with a p-value attached.

Common mistakes

  • Too many items. The most frequent error by a wide margin. A 15-item ranking generates a full data table where most of the middle is guesswork.

  • No ties allowed when respondents genuinely tie. Forcing an order between two items a respondent values equally manufactures a difference that does not exist. If ties matter to your decision, allow them or use a rating question alongside.

  • No escape hatch. If some items are irrelevant to some respondents, give them a way to say so. Without one, irrelevant items get parked at the bottom and become indistinguishable from actively rejected items.

  • Requiring a complete ranking. Mandatory full rankings drive abandonment and, worse, drive straight-lining from respondents who just want the screen to end. Optional or partial ranking gives you fewer cells and better data.

  • Mixing dissimilar items. Ranking only works when items are comparable on one dimension. Asking someone to rank "price", "customer support" and "the mobile app" together is asking them to rank a cost, a service and a product surface on an unstated criterion. Keep the list to one kind of thing and state the criterion explicitly in the question stem.

  • Vague criteria. "Rank these" is not a question. "Rank these in the order you would want us to build them" is.

Ranking tells you the order, not the reason

The structural limit of every ranking question is that it strips out reasoning. You learn that pricing came first. You do not learn whether that is because your prices are too high, because the pricing page is confusing, or because a competitor just undercut you. Those three findings lead to three completely different projects.

The usual fix is to follow the ranking with an open text box, which most respondents skip or answer in four words. A more productive approach is to probe conversationally: ask why the top item was ranked first, in the respondent's own words, immediately after they rank it. That is the reasoning ranking removes, and it is the part that tells you what to do.

This is where AI-moderated surveys change the economics. SmartInterview runs voice surveys with AI follow-up probing, so a ranking question can be followed by a spoken "why did you put that first?" and the open responses get coded automatically instead of sitting unread in a spreadsheet column. You keep the quantified priority order and recover the explanation behind it, which is the split that usually forces teams to choose between qualitative and quantitative research in the first place.

Frequently Asked Questions

How many items should a rank order question have?

For a full ranking, roughly 5 to 7 items is a practical working ceiling. This is a rule of thumb from survey design practice, not a hard research threshold. Beyond that, respondents reliably place their top and bottom items and effectively guess at the middle. For lists of 8 to 12, ask for the top 3 instead. Above about 15 items, switch to MaxDiff.

What is the difference between a rank order question and a rating scale?

A rating scale measures each item independently on an absolute scale, so every item can score highly. A rank order question forces a trade-off, so only one item can be first. Ratings tell you how much respondents value something and are comparable across people. Rankings tell you what comes first and are relative to each respondent's own list. Ratings are prone to everything scoring "very important"; rankings cannot tell you whether the top item matters at all.

Can you calculate an average of ranking data?

You can compute a mean rank, but read it carefully. Ranks are ordinal, so the mean assumes the distance between ranks 1 and 2 equals the distance between ranks 4 and 5, which the data never establishes. Use mean rank as a rough ordering device, always show the distribution of rank positions next to it, and use nonparametric tests such as Friedman or Wilcoxon rather than t-tests when testing differences.

Is drag-and-drop ranking bad for mobile surveys?

It is risky. On touch screens, dragging competes with page scrolling, long lists trigger awkward auto-scroll while an item is held, and mis-drops are easy. Drag-and-drop is also difficult for keyboard and screen reader users. If a meaningful share of your sample is on mobile, use numbered entry or a "tap your top 3 in order" format instead.

When should I use MaxDiff instead of ranking?

Use MaxDiff when the list is long, typically 15 items or more, and the goal is to find the leading handful. It shows small sets of items and asks for the best and worst in each, so respondent effort stays flat as the list grows. It also produces utility scores on a common scale, which supports comparisons ranking cannot make. The cost is more screens, an experimental design and a scoring model, so it is overkill for short lists.

Should respondents be forced to rank every item?

Usually not. Mandatory full rankings increase abandonment and encourage careless ordering just to get past the screen. Partial ranking collects the reliable part of the answer. Also give respondents a way to flag items that do not apply to them, otherwise irrelevant items get dumped at the bottom and look the same as items they actively reject.

Related articles

Sign up for free

Sign up for free

Sign up for free