Performance
Appraisals that calibrate across the whole team
Two managers sit in the same review cycle. One is generous - a 4 from them means "did the job". The other is exacting - a 4 means "genuinely excellent, rare". Their people are scored on the same 1-to-5 scale, and the scores are compared for promotions, pay and development. The comparison is meaningless, and everyone senses it, which is how appraisals quietly lose the trust of the people they are meant to develop.
Fixing this needs two things most appraisal templates skip: weighting the criteria, and calibrating the ratings as a set. Get both and the appraisal becomes something a promotion panel can actually rely on.
A flat average is a hidden decision
Score someone on five competencies and take the plain average, and you have silently declared that every competency matters equally. For most roles that is false. Technical quality might genuinely matter more than, say, willingness to volunteer for extra tasks - and a flat average buries that judgement instead of making it.
Weighting brings the decision into the open. Set each competency's weight deliberately - the important ones carry more - with one rule that keeps the model honest: the weights must total 100%. A live check on that total stops the classic error of weights that quietly sum to 95 or 110 and skew every score. Now the overall figure is a weighted score, and it reflects what the role actually values rather than an accidental equality.
From a number to a defensible rating
A weighted score to two decimal places is precise but not yet meaningful. Band it: a score below one threshold reads as "Needs improvement", the middle band as "Meets", the top band as "Exceeds", colour-coded so the rating is legible at a glance. The bands are where the number becomes a judgement a person can be given and a panel can act on.
The point of driving the rating from weighted competencies, rather than a manager simply picking an overall grade, is defensibility. When someone asks why a rating is what it is, the answer is not "that felt about right" - it is a visible line of scores against defined competencies, weighted by a model everyone agreed in advance. That is the difference between an appraisal that survives a challenge and one that does not.
Calibration: the step that makes ratings comparable
Here is the part that most templates leave out entirely. An individual rating means little in isolation; it only means something relative to the rest of the team. Calibration is the discipline of looking at the whole distribution before anything is signed off - how many people sit at each rating across every manager - and asking whether the shape is honest.
If every single person is rated "Exceeds", the ratings are not measuring performance, they are measuring managerial generosity. The distribution is the tool that catches that before it reaches a pay review.
You are not forcing a rigid curve onto people - forced ranking has real downsides and can poison a team. You are sense-checking. When one manager's team is all top-band and another's is all middle for similar work, that is a calibration conversation, not a promotion list. A native chart of how many people sit at each rating makes the shape impossible to ignore, so drift gets caught while it can still be corrected.
Running it well
- Agree the model before the cycle. Competencies, weights and band thresholds set in advance are fair; set afterwards, they look like they were reverse-engineered to justify a decision.
- Calibrate as a group. Managers reviewing each other's distributions is uncomfortable and exactly the point - it is the only reliable check on individual bias.
- Keep the framework editable. Roles differ, and the competencies and weights that fit engineering will not fit sales. A model you can adapt beats a fixed one you have to fight.
- Separate the rating from the pay outcome in the conversation. The appraisal is about growth; bolting the pay decision directly onto the same meeting makes people manage the score instead of the work.
One boundary to keep in view: performance ratings are personal data and are used to make decisions about people, so handle the records under your UK GDPR obligations and remember that a template structures the judgement - it does not replace it, and it is not a substitute for HR or legal advice on how you manage performance.
Weighted scoring, banded ratings and a calibration view in one workbook
Set competency weights to total 100%, score 1 to 5, and the weighted rating bands itself - then a calibration tab and distribution chart show the shape of the whole team before sign-off. A template to help you run a fair review, not HR or employment-law advice.