THE NEW DYNAMICS BLOG

Performance rating scale: 1 to 5 definitions, examples and how to choose

Get clear 1 to 5 rating scale definitions for performance reviews, compare 3, 4 and 5-point scales, see an anchored example and learn how to keep ratings fair.

Published Updated 11 min read

TL;DR

  • A performance rating scale turns a judgement about someone's work into a level, usually from 1 to 5. It is useful only when every level is defined in words that managers read the same way.
  • Level 3 should mean a fully successful year and should be the most common rating. Describe each level by what the person did, and anchor it with examples of behaviour for each job family.
  • Scales do not make ratings fair on their own. Evidence kept all year, behavioural anchors, calibration between managers and a conversation about the reasons matter more than whether you use three, four or five points.

Two managers give the same score of 3 out of 5. One means “solid, reliable and doing everything we asked”. The other means “a bit disappointing”. The employees compare notes, and both lose faith in the process. The problem is not the number. It is that nobody defined it.

This article gives plain 1 to 5 performance rating scale definitions, compares the common types of scale, shows a behaviourally anchored example and explains how to keep ratings consistent and fair.

If you want ratings to rest on goals, feedback and evidence gathered through the year, see how New Dynamics performance reviews bring them into one conversation.

What is a performance rating scale?

A performance rating scale is a set of defined levels that managers use to summarise how well someone has performed against what was expected. Most scales have between three and five levels, and each level has a label and a description.

The United States Office of Personnel Management describes rating as “evaluating employee or group performance against the elements and standards in an employee's performance plan, summarizing that performance, and assigning a rating of record”. It adds that organisations find it useful to summarise performance in this way to compare it “over time or among various employees”.

That is the case for a scale: it gives a common language, makes comparisons possible and supports decisions about pay, promotion and development. The case against is that a single number can hide more than it reveals. We return to that below.

1 to 5 rating scale definitions

The definitions below are ours. Adapt the wording, but keep two principles: describe what the person did, and make level 3 a good result.

LevelLabelDefinition
5OutstandingResults far beyond what was agreed, sustained through the year. The person's work raised the standard for others. Rare by definition.
4Exceeds expectationsMet every goal and clearly went beyond several. Needs little direction, and makes colleagues more effective.
3Fully meets expectationsAchieved what was agreed, to the standard required, and showed the behaviours expected. A successful year.
2Partly meets expectationsMet some goals and missed others, or met them with more support than the role should need. Specific improvement is required.
1UnsatisfactoryDid not meet most goals or the basic standards of the role, despite support. Formal improvement action is needed.

Why level 3 has to be a good rating

In a healthy organisation, most people do what was asked of them and do it well. They belong at level 3. If a 3 feels like a consolation prize, managers will inflate ratings to avoid hurting people, and the scale drifts upwards until a 4 is the norm.

Avoid the label “average” or “satisfactory” for level 3. “Fully meets expectations”, “successful” or “strong performer” says what you mean.

What each level might look like

For a customer service adviser, with illustrative details:

  • 5: resolved the most complex cases in the team, wrote the guide that others now use and kept satisfaction above 4.8 out of 5 all year.
  • 4: consistently beat the targets for response time and satisfaction, and coached two new starters.
  • 3: met the targets for response time and satisfaction, handled difficult customers courteously and was reliable throughout.
  • 2: met the response target but fell short on satisfaction in two quarters. Needed frequent help with standard cases.
  • 1: missed most targets, and errors continued after training and support.
A 1 to 5 performance rating scale with definitions: 1 unsatisfactory, 2 partly meets expectations, 3 fully meets expectations, 4 exceeds expectations and 5 outstanding.
A number means nothing until you define it in words.

Types of rating scale

ScaleExample levelsStrengthsWeaknesses
Three-pointBelow, meets, exceedsSimple, quick and hard to argue overLittle room to recognise the truly exceptional
Four-pointUnsatisfactory, developing, successful, exceptionalNo middle option, so managers must decideCan feel forced, and “developing” becomes a dumping ground
Five-pointAs in the table aboveFamiliar, with room at both endsInvites clustering on the middle score
Behaviourally anchoredEach level described by specific behaviours for the jobConsistent between managers, and tells people what to changeTakes time to write for each job family
No overall ratingNarrative assessment, with a separate pay decisionKeeps attention on the conversationPay decisions still need a basis, which can become hidden

The textbook Organizational Behavior, published by OpenStax, notes in its section on techniques of performance appraisal that one of the most serious drawbacks of the simple graphic rating scale is “its openness to central tendency, strictness, and leniency errors”. In other words, the scale is only as good as the definitions behind it.

Behaviourally anchored rating scales

A behaviourally anchored rating scale, or BARS, describes each level with specific behaviours. OpenStax explains the benefit: raters are considering “verbal descriptions of specific behaviors instead of general categories of behaviors”, which should reduce many rating errors.

Here is an anchored scale for one behaviour, collaboration.

LevelAnchor
5Builds cooperation between teams. Others seek this person out to resolve conflicts, and joint work improves because of them.
4Shares information before being asked, offers help to colleagues under pressure and gives credit publicly.
3Works well with colleagues, shares what others need to know and raises disagreements directly and politely.
2Cooperates when asked, but often works alone. Sometimes holds back information that others need.
1Undermines colleagues, withholds information or creates conflict that damages the team's work.

Compare that with a scale that says only “Collaboration: 1 2 3 4 5”. With anchors, a manager has to find the description that fits the evidence. The employee can see what the next level requires.

Our guide to core competencies shows how to write behaviours at several levels.

From a bare rating, 'Collaboration: 3 out of 5', to an anchored level: 'Shares what others need to know and raises disagreements directly and politely.'
OpenStax: raters consider “verbal descriptions of specific behaviors instead of general categories”.

What to rate

Goals. Rate achievement against each goal agreed in the performance plan: not met, partly met, met or exceeded. Allow for factors outside the person's control.

Behaviours. Rate a short list of behaviours with anchored descriptions.

An overall rating, if you use one. Base it on both, and on judgement. Avoid averaging scores to a decimal place. A rating of 3.47 claims a precision that does not exist.

Someone who hits every target while damaging the team has not had an outstanding year. Make sure the scale allows you to say so.

How to keep ratings fair

OpenStax lists the classic problems with performance appraisals, including central tendency, which it calls “the failure to recognize either very good or very poor performers”, strictness or leniency, the halo effect, recency and personal bias. A scale does not remove them. These practices help.

  • Define every level in writing, with examples for each job family.
  • Keep evidence through the year, so that the rating covers twelve months and not the last six weeks.
  • Ask for a self-assessment on the same scale. See our self-evaluation examples.
  • Calibrate. Managers compare proposed ratings and the evidence behind them before anything is final. Our performance calibration guide explains how.
  • Check the patterns. Look at ratings by team, grade, gender, ethnicity, age and working pattern. Unexplained gaps need attention. See our unconscious bias examples.
  • Do not force a curve. Requiring fixed percentages at each level guarantees that some good performers are labelled poor. The CIPD performance management factsheet describes a move towards “less focus on process, such as forced ranking or guided distribution ratings”.
  • Explain the rating. The person should hear the reasons and the evidence, and have the chance to respond. Our performance review phrases can help with wording.

Should you drop ratings altogether?

Some organisations have removed overall ratings. The argument is that people fixate on the number and stop listening to the feedback.

There is something in that. But pay and promotion decisions still get made, and without a visible rating the basis for them can become less transparent. If you drop ratings, be clear about how those decisions will be reached. We discuss the trade-offs in performance reviews: the good, the bad and the ugly.

A middle path is common: rate goals and behaviours separately, hold the development conversation first and discuss pay at a different time.

How to build a rating scale in six steps

  1. Decide what the scale is for. Development, pay, promotion or all three. The purpose affects how many levels you need.
  2. Choose the number of levels. Use the fewest that serve the purpose. Three or four is enough for many organisations.
  3. Write the labels and definitions. Describe what the person did. Make the middle level a positive result.
  4. Add behavioural anchors for your main job families, with managers and employees involved in the wording.
  5. Test it. Ask several managers to rate the same written case studies. Where they disagree, fix the wording.
  6. Train, calibrate and review. Brief managers before each cycle, calibrate every time and look at the distribution of ratings each year.
Six steps to build a performance rating scale: decide the purpose, choose the number of levels, write labels and definitions, add behavioural anchors, test with case studies, then train, calibrate and review.
Where managers disagree on a test case, fix the wording.

Common mistakes

Numbers without definitions. The opening example of this article.

A middle level that feels like failure. It drives inflation.

Too many levels. On a ten-point scale nobody can explain the difference between a 6 and a 7.

False precision. Weighted averages to two decimal places.

Surprises. A low rating should never be the first time that someone hears about a problem.

Rating the person. “She is a 2” is a label. “This year's results partly met expectations, for these reasons” is an assessment.

Setting and forgetting. Definitions drift. Review the scale and the distribution of ratings every year.

Frequently asked questions

What do the numbers on a 1 to 5 rating scale mean?

A common set of definitions is: 1 unsatisfactory, 2 partly meets expectations, 3 fully meets expectations, 4 exceeds expectations and 5 outstanding. Each organisation should define its own levels in writing, with examples, so that managers interpret them in the same way.

Is a 3 out of 5 a good performance rating?

It should be. On a well-defined scale, 3 means that the person achieved what was agreed to the standard required, which is a successful year. If a 3 feels like a poor result in your organisation, the labels or the culture around ratings need attention.

What is the best rating scale for performance reviews?

There is no single best scale. Three and four-point scales are simple and reduce hair-splitting. Five-point scales give more room at the ends. Whatever you choose, written definitions, behavioural anchors and calibration matter more than the number of points.

What is a behaviourally anchored rating scale?

A behaviourally anchored rating scale, or BARS, describes each level with specific examples of behaviour for the job. OpenStax notes that raters then consider descriptions of specific behaviours, not general categories, which should reduce rating errors.

Should you use a forced distribution for ratings?

We advise against it. Forcing fixed percentages into each level guarantees that some good performers receive poor ratings, and it sets colleagues against each other. Use calibration between managers to keep standards consistent.

How do you make performance ratings fair?

Define each level in writing, collect evidence all year, ask for a self-assessment, calibrate ratings between managers, check the distribution for unexplained gaps between groups, and explain every rating to the employee with reasons and evidence.

Your next step: test your definitions

  1. Write down what a 3 means in your organisation, in one sentence.
  2. Ask three managers to do the same, without conferring.
  3. Compare the answers.
  4. If they differ, agree the definition and one example before the next review cycle.

For a step-by-step way to compare ratings and evidence between managers, read and download our performance calibration guide. The guide is free to read, and the PDF uses our short download form.

Want ratings that rest on goals, feedback and evidence from the whole year? Book a New Dynamics demo and bring your current scale. You can also email contact@new-dynamics.com.

Keep the conversation going.

Bring out the best
in your people.

See what performance management could look like for your organisation.

Book a demo