QA calibration is a structured session where two or more evaluators independently score the same recorded call, then compare scores to identify and resolve interpretation gaps before applying the scorecard to production evaluations. The goal is to ensure that a given call receives the same score regardless of which evaluator reviews it.
A QA scorecard is only as useful as the people scoring it are consistent. If two analysts evaluate the same call and produce scores fifteen points apart, neither score tells you much about the agent — they tell you about the evaluators. Calibration is the process that closes that gap, aligning reviewers on what each criterion actually means so that scores become a reliable measure of agent performance rather than a proxy for evaluator opinion.
Why calibration is necessary
Without calibration, QA scores drift. Two evaluators reading the same scorecard will form slightly different mental models of what "active listening" or "empathy" means in practice. One might weight tone heavily; another might focus on whether the agent paraphrased the customer's concern. Over time these interpretive differences compound: one team lead's agents appear to score higher not because they perform better but because their evaluator grades more leniently.
This evaluator bias makes QA scores meaningless as coaching tools. An agent who receives a 74 from one analyst and a 84 from another on the same call cannot draw any useful conclusion from either number. The score reflects the reviewer's subjective interpretation as much as anything the agent did. When agents notice that scores vary based on who is reviewing, trust in the QA program erodes — and with it, the program's influence on behavior.
The statistical objective of calibration is inter-rater reliability (IRR) — the degree to which independent raters agree when evaluating the same subject. In contact center QA, IRR is typically measured as the percentage of evaluation items where two or more reviewers agree within an acceptable tolerance. High IRR means scores can be compared across agents, teams, and time periods with confidence. Low IRR means the data is contaminated by evaluator variance and should not be used for performance management decisions.
How to run a calibration session
Step 1: Select a calibration call
Choose a recorded call that exercises contentious or ambiguous scorecard criteria — calls where reasonable evaluators might disagree. A straightforward call where every criterion is clearly met or clearly missed teaches evaluators nothing. Look for calls that include borderline moments: an agent who de-escalated imperfectly, a compliance disclosure delivered but rushed, or empathy language that could be read as genuine or scripted. Rotate the types of calls used across calibration sessions to cover different products, call reasons, agent skill levels, and customer profiles. This prevents evaluators from becoming calibrated only to one narrow call type.
Step 2: Independent scoring
Each evaluator scores the selected call separately, without seeing or discussing other evaluators' scores. This independence is essential: if evaluators confer before scoring, the session measures social dynamics rather than individual interpretation. Evaluators should complete the full scorecard and note their reasoning for any items they find ambiguous, because those notes will be the raw material of the discussion phase.
Step 3: Score comparison and gap identification
Reveal scores simultaneously — either in the calibration platform, a shared spreadsheet, or verbally if the group is small. For each scored item, flag cases where scores differ by more than the tolerance threshold. A common starting threshold is ±10 points on a 100-point scale, but many teams use tighter tolerances as their programs mature. Items within tolerance are noted and moved past quickly. Items outside tolerance become the agenda for the discussion phase. Do not skip items just because the overall call score happens to be similar — two evaluators can arrive at the same total through different routes, masking disagreement that will surface on the next call.
Step 4: Discussion and resolution
For each flagged gap, each evaluator explains their reasoning. The goal is not to determine who was right and who was wrong — it is to understand what interpretation each evaluator was applying and to agree on a shared standard going forward. Someone on the call (typically the QA program lead) documents the agreed interpretation and any illustrative language that helped the group align. This agreed interpretation becomes a scoring note attached to that criterion in the rubric.
When evaluators genuinely cannot resolve a gap, escalate to a senior QA manager or program owner to make the determination. Unresolved gaps left open will resurface in every subsequent session involving the same criterion.
Step 5: Update the calibration log
Record the session in a calibration log: which call was used, who attended, each evaluator's score, the items that produced gaps, and the resolution for each gap. This log is the program's institutional memory. When a new evaluator joins the team, the calibration log — especially its resolution notes — is far more useful than the scorecard alone. The log also allows the program to track whether the same criterion keeps producing disagreement, which is a signal that the criterion itself needs to be rewritten.
Step 6: Apply the learning
If the discussion surfaces an ambiguous criterion — one where the scorecard language genuinely supports more than one reasonable interpretation — update the scorecard rubric or scoring guide immediately rather than waiting for the next revision cycle. Ambiguous criteria do not become clearer through additional calibration sessions; they continue generating the same disagreements. Rewrite the criterion with specific, observable behavioral examples of what a full score, a partial score, and a zero score look like.
Tolerance thresholds and acceptable variance
Not every scoring gap requires the same response. A structured tolerance framework helps teams spend calibration time on genuine disagreements rather than minor rounding differences.
A widely used framework:
- ±5 points or fewer: Acceptable variance. Evaluators are essentially aligned. Note the item and move on.
- 6–10 points: Discussion required. Evaluators explain their reasoning. If a clear interpretation emerges, document it and close the gap.
- More than 10 points: Mandatory resolution before calibration is considered complete. The gap is significant enough that leaving it unresolved will produce meaningfully different scores in production evaluations.
Auto-fail criteria operate outside this framework entirely. Auto-fail items — typically compliance disclosures, identity verification steps, required legal language, or safety protocol adherence — must score identically across all evaluators. There is no acceptable variance on a criterion that fails the entire call when missed. If evaluators disagree on whether an auto-fail criterion was met, that disagreement must be resolved and the criterion's definition must be tightened until identical calls produce identical auto-fail determinations.
Evaluator bias types to watch for
Calibration sessions are also the primary mechanism for identifying systematic evaluator bias. The most common patterns:
- Leniency bias: An evaluator consistently scores higher than peers across multiple calls and multiple criteria. Often reflects discomfort with assigning low scores or an overcorrection from wanting to protect agent morale. Leniency bias inflates scores and masks genuine performance gaps.
- Severity bias: The inverse — an evaluator consistently scores lower than peers. May reflect high standards, but systematic severity bias produces agents who feel unfairly penalized regardless of improvement. If severity bias is individual rather than program-wide, the evaluator's scores are not comparable to others'.
- Recency bias: Evaluators weight the end of a call disproportionately when completing the scorecard. A call that ended positively may receive inflated scores on earlier criteria; a call that ended badly may receive deflated scores on earlier strong performance. Evaluators should score each criterion at the moment it occurs or immediately after, rather than completing the scorecard after the call ends.
- Halo effect: A single strong behavior early in the call causes the evaluator to score subsequent criteria more generously than the evidence warrants. The inverse — the horn effect — occurs when a single early weakness depresses subsequent scores. Structured scorecards with criterion-by-criterion scoring reduce but do not eliminate halo and horn effects.
When calibration reveals that one evaluator's scores systematically diverge from the group in one direction, the appropriate response is not to simply average the scores — it is to investigate whether bias is present and, if so, to provide targeted coaching to that evaluator on the affected criteria.
Calibration frequency and who should attend
Active QA programs typically run calibration sessions weekly or bi-weekly. Monthly calibration is a minimum for any program that uses QA scores to inform performance management decisions. Programs that run calibration less frequently than monthly tend to accumulate interpretive drift between sessions that is difficult to correct retroactively.
Core attendees should include all QA analysts who complete production evaluations. Team leads who use QA scores to coach their agents should attend regularly — ideally every session — because their coaching is downstream of evaluation decisions, and their understanding of the scoring rationale matters. Agents themselves can attend calibration sessions in a listening-only capacity, which builds transparency and helps agents understand precisely how scoring decisions are made. Agents who sit in on calibration often report that the experience is more instructive than any post-evaluation debrief, because they hear the full deliberation rather than just the outcome.
Calibration sessions should be kept separate from production evaluation sessions. If evaluators use the calibration call as a production evaluation for the featured agent, it creates a conflict of interest — evaluators may anchor their calibration score to the group average rather than their independent judgment, contaminating both the calibration and the production record.
Maintaining calibration over time
A QA program that calibrated successfully in its first month does not remain calibrated indefinitely. Evaluator turnover, new product launches, changes to the scorecard, and shifts in call volume all introduce new sources of interpretive variance. Treat calibration as a continuous process rather than a one-time setup task.
Trigger a re-calibration cycle when:
- New criteria are added to the scorecard — new language requires new shared interpretation
- Score averages shift unexpectedly across the program — an unexplained rise or fall in average QA scores often reflects evaluator drift rather than real performance change
- Agent or team lead complaints about scoring inconsistency increase — agents notice evaluator variance before management does
- A new evaluator joins the program — onboarding calibration should precede any production evaluation responsibility
Track inter-rater agreement scores over time. Plot the average gap between evaluators for each session and each criterion. Criteria that consistently produce high inter-rater variance are candidates for rubric revision. Evaluators whose individual scores consistently diverge from the group median are candidates for targeted coaching. A QA program that monitors its own reliability data continuously improves faster than one that treats calibration as a periodic compliance exercise.
Frequently Asked Questions
Call recordings are the raw material of every calibration session — see What Is Call Recording? for how recordings are captured and stored. Live monitoring modes that generate real-time coaching data are covered in Listen, Whisper, Barge, and Intercept. For the broader metrics framework that QA scores feed into, see Contact Center Analytics.