Back

Customer Satisfaction Metrics: CSAT, NPS and CES Compared

SmartInterview Team

The Short Answer

Three customer satisfaction metrics dominate. CSAT measures how happy someone was with a specific interaction. NPS measures willingness to recommend the brand. CES measures how hard the customer had to work. They are not interchangeable, and none of them tells you what to fix. The open follow-up question does that.

  • CSAT: "How satisfied were you?" on a 1-5 scale. Reported as the percentage choosing the top one or two boxes.

  • NPS: "How likely are you to recommend us?" on 0-10. Reported as %promoters minus %detractors, from -100 to +100.

  • CES: "How easy was it to get this done?" usually on 1-7. Reported as an average or a top-box percentage.

  • Pick one as your anchor and run it properly. Three metrics run badly is worse than one run well, because you get three unreliable trend lines instead of one you believe.

  • Your only safe benchmark is your own history, measured the same way on the same population.

Working with NPS specifically? Start with our Net Promoter Score guide or run the numbers in the NPS calculator.

CSAT vs NPS vs CES

The single table worth keeping. Each metric is good at one thing and blind to something else.

Metric

Question asked

Scale

Best for

Blind spot

CSAT

How satisfied were you with [this interaction]?

Usually 1-5, reported as % top-box

Rating one specific moment: a delivery, a call, an order

Skews high and ceilings out; says nothing about loyalty

NPS

How likely are you to recommend us to a friend or colleague?

0-10, reported -100 to +100

Tracking overall relationship health over time

Discards most of the scale; identical scores hide opposite distributions

CES

How easy was it to get your issue handled?

Usually 1-7

Diagnosing friction in service and support journeys

Silent on affection, price and product quality

Open follow-up

Why did you give that score?

Free text or voice

Explaining any of the three above

Needs coding before it scales past a few hundred responses

Which one to run

If you want to know…

Use

Did that support call go well?

CSAT, or CES if the issue was procedural

Is the relationship getting stronger or weaker?

NPS on a fixed cycle

Where is our service process painful?

CES, per journey step

Why is the number moving?

The open follow-up, coded

How Each Metric Is Calculated

CSAT

Ask "how satisfied were you?" on a labelled scale, most commonly five points from very dissatisfied to very satisfied. Then report the percentage of respondents who chose the top one or two boxes.

CSAT = (respondents choosing the top boxes / total respondents) x 100.

With 200 responses where 120 said "very satisfied" and 40 said "satisfied", top-two-box CSAT is (160 / 200) x 100 = 80%.

The thing to nail down and never change is which boxes count. Top-one-box and top-two-box produce very different numbers from identical data, and quietly switching between them is a common way to manufacture an improvement. Write the definition into the reporting spec.

NPS

Ask likelihood to recommend on 0-10. Promoters score 9-10, passives 7-8, detractors 0-6. Subtract the percentage of detractors from the percentage of promoters.

NPS = %promoters minus %detractors.

Two traps. Passives stay in the denominator: you divide by everyone who answered, not just promoters plus detractors. And the result is written as 25 or +25, never 25%, because a difference between two percentages is not itself a percentage.

CES

Ask how much effort the customer had to expend, usually as agreement with a statement like "the company made it easy for me to handle my issue", on a 1-7 scale. Report either the mean or the percentage giving the top two or three ratings.

Direction matters here and it catches people out. Some implementations ask about ease, where higher is better. Others ask about effort, where higher is worse. Mixing the two across teams produces a metric that means the opposite of what half the room assumes. Pick the ease framing, write it down, and never invert it.

What Each Metric Cannot Tell You

CSAT cannot tell you about loyalty

Satisfaction with an interaction is a weak predictor of whether someone stays. Customers routinely rate an interaction highly and leave anyway, because the interaction was fine and the product, the price or the alternative was better. CSAT also runs into a ceiling problem: satisfaction questions attract generous answers, and once you are reporting in the high eighties, real degradation can happen inside your reported range without visibly moving it.

NPS cannot tell you about distribution

Because only the net matters, wildly different customer bases produce identical scores. 40% promoters and 40% detractors gives 0. So does 100% passives. Those are two different businesses with two different action plans. Always publish the split alongside the score.

NPS is also noisier than people expect. It is a difference between two proportions, so its sampling error is wider than a single proportion at the same base. A lot of the month-to-month movement on NPS dashboards is nothing but small-sample noise being narrated as a trend.

CES cannot tell you whether they love you

Effort is a strong lens on service processes and a poor one on everything else. A frictionless experience of a product nobody wants scores well on CES and predicts nothing. CES belongs on service journeys, not as a company-level headline metric.

None of them tells you what to change

This is the shared limitation and it is the important one. All three compress an experience into a number. Numbers tell you the direction and roughly where to look. None of them names the failure, and no amount of extra scales will change that, because a closed scale can only report on things you thought to ask about when you wrote it.

Why Running All Three Badly Is Worse Than One Well

The instinct when a measurement program feels weak is to add metrics. It usually makes things worse, for four reasons.

Each metric splits your sample

You have a finite number of customers willing to answer anything. Splitting them across three instruments gives you three bases too small to report on confidently, each swinging on the arrival of a handful of responses. One metric on the full sample gives a trend line you can defend.

Three metrics means three arguments

When CSAT rises, NPS falls and CES is flat, meetings become debates about which number to believe. In practice the team quietly settles on whichever one is moving in a favorable direction, and the measurement program stops functioning as a check on anything.

Survey fatigue is cumulative

Customers do not distinguish between your CSAT program and your NPS program. They experience three surveys. Each additional ask lowers the response rate on all of them, and it does not lower it randomly: the people who disengage first are usually the ones who had already told you something and seen no result.

Nobody has time to act on three sets of verbatims

If you were only going to route and act on one stream of open responses, running three collects two streams that go nowhere. That is worse than not asking, because you have now taught customers that answering you does nothing.

A defensible setup

  1. One relationship metric, run on a fixed cycle to a cross-section of customers. NPS if your organization already reports it, since breaking a trend line has a real cost.

  2. One transactional metric, triggered by events. CSAT for general interactions, CES where the journey is procedural.

  3. One open follow-up on both, coded into a stable theme frame.

  4. Nothing else until those three are running reliably and someone is acting on the themes.

Keep the relational and transactional lines separate and label them. They sample different people and answer different questions, so combining them into one company number is a category error. Support-ticket scores in particular over-represent customers who had something go wrong.

Benchmarking Traps

External benchmarks are the most requested and least reliable part of a satisfaction program. Four traps recur.

Category effects dominate

Some categories can generate enthusiasm and some structurally cannot. People recommend a restaurant or a favorite app readily. Almost nobody spontaneously recommends their electricity supplier or their insurance claims process, however competently those run. Comparing across categories tells you about the category, not about performance.

Country and language shift the scores

Cross-cultural survey research has repeatedly documented that respondents in different countries use rating scales differently, with some populations reaching for the extremes more readily than others. Because NPS counts only 9s and 10s, a market that habitually avoids the top of a scale will produce a structurally lower score for identical underlying sentiment. If you run the same survey in several markets, expect part of the ranking between them to be an artifact of scale use rather than a real difference. Compare each market against its own history.

Method differences are invisible in the headline

A benchmark collected by phone, from a general population panel, on a relationship question is not comparable with your number collected by email, from recent purchasers, right after a transaction. The two figures share a name and measure different things. Before using any published benchmark, check the mode, the sample and the trigger. If those are not disclosed, the benchmark is not usable.

Self-selection inflates published figures

Companies publish good scores and stay quiet about bad ones, so any informal collection of "industry averages" scraped from vendor marketing is biased upward. Treat unsourced benchmark numbers as marketing.

The practical rule: your own score, measured the same way on the same population, over time, is the only benchmark you can fully trust. If you need an external comparison, buy a syndicated study that publishes its methodology and check that it collected data the way you do.

The Follow-Up Question That Makes Any Metric Actionable

If you can only improve one thing about your satisfaction program, improve the question after the score. The score gives you a direction. The follow-up gives you something to do on Monday.

Anchor the follow-up to the score they just gave

"Why did you give that score?" pulls thin answers like "good service". Questions that reference the specific rating pull specifics:

  • To low scorers: "What went wrong, and what would have needed to happen for this to go well?"

  • To middle scorers: "What is the one thing that would have made this a 10?"

  • To high scorers: "What specifically do you tell other people about us?"

The high-scorer question is the underused one. It hands you your own positioning language, written by people who already buy.

Probe once, then stop

Nearly every first open answer stops one layer above the cause. "Delivery was slow" is a symptom. One targeted follow-up, asked while the person is still in the survey, gets you to the specifics: which order, what was promised, what they expected instead.

This is the layer a static form cannot deliver, and it is where AI probing earns its place. A model that reads the first answer and asks one relevant second question turns a fragment into a usable account without a moderator on the call. Voice pushes it further still, because people say considerably more out loud than they will type into a box, particularly on a phone. The mechanics are in voice surveys vs traditional surveys.

Code the answers or the program dies

Open text only survives contact with a business review once it is quantified. You need to be able to say "shipping delays account for 31% of low-scorer comments this quarter, up from 18%". That means every comment gets a theme code, the frame stays stable between waves so the trend means something, and you can click a theme to read the raw verbatims behind it. Manual coding does this well and stops scaling somewhere in the low hundreds per wave. Automated coding scales, but only counts if you can audit it against the original words.

Route it, then close the loop

The strongest driver of value in a satisfaction program is not the metric you chose. It is whether an unhappy customer hears back from a human within a working day or two, and whether the rest of your customers ever learn what changed. Score-only programs stall because nothing happens after the survey. If you want the operating model around that, read the customer feedback loop.

Where to Go Next

A score is a symptom. Get the cause.

Every satisfaction metric ends at the same place: a number that tells you something moved, and no explanation of why. The explanation is sitting in the open responses, which is exactly the part most programs collect and never read.

SmartInterview asks the score and then asks the question you would have asked in person. Respondents answer by voice or text, an AI probes the reason in their own language, and the open responses come back coded into themes you can count, trend and trace back to the exact words a customer used.

Keep the metric you already report and add the layer underneath it. Start free or book a demo.

Frequently Asked Questions

What are the main customer satisfaction metrics?

CSAT, NPS and CES. CSAT measures satisfaction with a specific interaction on a short labelled scale. NPS measures likelihood to recommend the brand on 0-10 and is reported from -100 to +100. CES measures how much effort the customer had to expend, usually on a 1-7 scale. Each is paired with an open follow-up question to be useful.

What is the difference between CSAT and NPS?

CSAT is transactional and asks about one experience; it is best for judging whether a specific process is working. NPS is relational and asks about the brand as a whole; it is best for a long-run trend line. They routinely move in different directions for the same company, which is expected rather than a sign that one is wrong.

How do you calculate a satisfaction score?

CSAT is the percentage of respondents choosing the top one or two boxes of the scale, so 160 top-box answers out of 200 is 80%. NPS is the percentage of promoters minus the percentage of detractors, with passives left in the denominator. CES is reported as a mean or a top-box percentage. In every case, write down the definition and never quietly change it.

Which customer satisfaction metric should I use?

One relationship metric on a fixed schedule, one transactional metric triggered by events, and an open follow-up on both. Use CSAT for general interactions, CES where the journey is procedural like support or returns, and NPS as the relationship anchor if your organization already reports it.

What is a good satisfaction score?

There is no defensible universal threshold. What counts as strong depends on the category, the countries your respondents live in, and how the data was collected, since scale-use habits and survey mode both shift the number. Compare against your own history, measured the same way on the same population.

Can I run CSAT, NPS and CES at the same time?

You can, but running all three usually leaves you with three bases too small to trust, three competing narratives in the meeting, and more survey fatigue than any of them is worth. One well-run metric with a coded open follow-up beats three shallow ones.

Why is my satisfaction score high but customers still leave?

Because satisfaction with an interaction is a weak predictor of loyalty. A customer can rate a support call five out of five and cancel the same month, because the call was fine and the product, the price or the competitor was the issue. Track a relationship metric alongside the transactional one, and read the open responses from the people who left.

Related articles

Sign up for free

Sign up for free

Sign up for free