Skip to content
CCaaS & Contact Center 9 min read

How to Build a Call Center QA Scorecard

QA scorecard layout with weighted evaluation sections, observable criteria, and auto-fail compliance gates

A QA program without a scorecard is just opinions. Different evaluators notice different things, weight them differently, and deliver feedback that agents cannot act on consistently. The scorecard is the instrument that converts a subjective listening experience into a structured, repeatable measurement — one that means the same thing whether a supervisor in one city evaluates a call or an evaluator in another city evaluates a different call the same morning.

Building that instrument well requires more than listing behaviors you want to see. It requires deciding which behaviors matter most, how to describe them so that any evaluator would score them the same way, which violations are severe enough to fail an entire call, and how to make the scorecard function as a coaching tool rather than just a report card.

A call center QA scorecard is a structured evaluation instrument that breaks agent performance into scored sections and criteria, enabling consistent, quantitative assessment of call quality across a team. Each section groups related behaviors, each criterion is scored against a defined standard, and the weighted total produces a call score that can be trended, benchmarked, and used to direct coaching.

Scorecard sections and their purpose

Most QA scorecards organize criteria into five or six sections that map to the natural arc of a customer interaction. Grouping related criteria serves two purposes: it makes scoring easier to follow, and it gives evaluators and agents a clear picture of which phase of the call had problems. A score breakdown by section is far more actionable than a single total.

Opening and greeting

The opening section covers the first twenty to thirty seconds of the call: did the agent answer within the required number of rings, introduce themselves and the company correctly, confirm they are speaking with the right person, and set a tone that is warm and appropriately paced? A poor opening can undermine the entire interaction before the customer has stated a single word about their issue.

Criteria in this section typically evaluate identity verification at the point of greeting (where required), use of the customer's name, and absence of behaviors that signal disengagement — rushed delivery, mispronounced company name, or failure to acknowledge the caller.

Needs identification and discovery

This section evaluates whether the agent asked enough of the right questions to fully understand what the customer needed before moving to resolution. Active listening markers — paraphrasing, acknowledgment statements, open-ended probing questions — belong here. So does the agent's ability to distinguish between the customer's stated request and their underlying need, which are not always the same thing.

Agents who skip this phase and jump to resolution based on incomplete information generate repeat contacts and lower first-call resolution rates. The scorecard makes that visible by scoring discovery as its own section rather than folding it into resolution.

Resolution and product knowledge

Resolution criteria assess whether the agent provided an accurate, complete answer to the customer's issue. This is typically the highest-weighted non-compliance section because it maps most directly to the reason the customer called. Criteria include: was the information provided factually correct, was the resolution complete or partial, did the agent take appropriate ownership rather than deflecting, and was the first-call resolution achieved or was a follow-up necessary?

Product knowledge is evaluated here as well. An agent who provides a confident but incorrect answer scores worse than one who acknowledges uncertainty and escalates appropriately, because the former creates downstream problems the customer will have to call back about.

Compliance and regulatory

Compliance criteria cover legally or contractually required behaviors: reading required disclosures, completing mandated verification steps, following script requirements where they apply, and avoiding prohibited statements. In regulated industries — financial services, healthcare, collections, insurance — this section carries the most weight and is most likely to contain auto-fail items.

Even in less regulated environments, compliance criteria often include internal policies: not making unauthorized commitments, not discussing competitor pricing in certain ways, not offering discounts beyond agent authority. These protect the organization even when no external regulator is watching.

Call handling mechanics

This section covers the procedural elements of managing a call: placing the customer on hold correctly (announcing the hold, giving a time estimate, checking back if the hold runs long), executing a transfer with the appropriate warm handoff or cold-transfer procedure, completing after-call work within the target window, and logging the call accurately. These behaviors are largely binary — either the agent followed the procedure or did not — which makes them among the easiest to score consistently.

Closing

The closing section evaluates whether the agent brought the call to a proper conclusion: summarizing what was resolved or agreed, confirming the customer's next steps, asking whether anything else is needed, and ending the call with appropriate acknowledgment. A weak closing leaves customers uncertain about what happens next and contributes to repeat contacts. Evaluators look for whether the agent owned the close or let the call trail off.

Weighting — not all criteria are equal

Once you have your sections and criteria, you need to assign weights. Weighting reflects your organization's priorities: a criteria set where compliance carries the same weight as hold procedure sends the wrong signal about what matters.

A common starting framework — illustrative only, to be adjusted to your business context — distributes section weights approximately as follows:

  • Compliance and regulatory: 30% — highest weight because failures here carry legal, financial, or reputational consequences
  • Resolution and product knowledge: 35% — weighted above communication because solving the customer's issue is the core purpose of the interaction
  • Communication and soft skills: 20% — important for customer experience but secondary to whether the issue was actually resolved correctly
  • Call handling mechanics: 15% — procedural adherence matters but rarely drives CSAT by itself

Within sections, individual criteria can carry different weights. A single item covering mandatory regulatory disclosure might represent 20 of the 30 compliance points, while a greeting-format criterion represents 2 of the opening's 10 points. Weight by business outcome: what does a failure on this criterion actually cost the customer, the organization, or the regulatory relationship?

Resist the temptation to weight every criterion equally because it is simpler. Equal weighting implies that missing the company's greeting script is as important as failing to verify caller identity before discussing account details. It is not, and a flat weighting system will not surface that distinction in your scores.

Observable versus subjective criteria

The single most common reason QA programs fail to produce consistent scores is criteria written around impressions rather than observable behaviors. The difference determines whether two evaluators listening to the same call would score it the same way.

A subjective criterion sounds like: "Agent was professional and courteous." There is no way to score this without applying personal interpretation. One evaluator's "professional" is another's "cold." The result is evaluator disagreement, which undermines calibration and makes the scorecard impossible to defend when an agent contests a score.

An observable criterion sounds like: "Agent verified caller name and account number before discussing account details." Either the agent did this or did not. The evaluator does not need to interpret anything. Observable criteria describe specific, audible behaviors that any evaluator can identify by listening to the recording.

Converting subjective criteria to observable ones requires asking: what would I actually hear if the agent did this correctly? Write that behavior down in specific terms. "Agent used empathy statement after customer expressed frustration" is more observable than "Agent demonstrated empathy." "Agent read the required fraud disclosure verbatim before processing the transaction" is more observable than "Agent followed compliance procedures."

Observable criteria reduce evaluator disagreement, make calibration sessions shorter and more productive, and give agents unambiguous feedback they can act on.

Auto-fail criteria

Auto-fail criteria are behaviors severe enough that a single instance fails the entire call, regardless of how well the agent performed on every other criterion. Auto-fails exist because some behaviors cannot be partially offset by good performance elsewhere. A call on which the agent misrepresented a product feature is not an 80-point call because the hold procedure was excellent.

Common auto-fail categories across industries:

  • Failure to verify caller identity before discussing or modifying account information — a direct security and compliance failure
  • Making a prohibited promise — committing the organization to something outside agent authority or policy (a specific refund amount, a callback on a date the system cannot support)
  • Missing a legally required disclosure — in regulated industries, omitting a mandated disclosure can carry regulatory consequences independent of whether the customer noticed
  • Misrepresenting a product or service — providing materially false information about features, pricing, or terms
  • Unprofessional conduct — abusive, discriminatory, or retaliatory behavior toward the customer

Auto-fail criteria must be documented explicitly, separately from weighted scoring criteria, and reviewed with all agents as part of onboarding and calibration. Every agent should be able to name your auto-fail behaviors without looking them up. When an auto-fail occurs, the call receives a zero or a designated failing score, and a separate coaching and documentation workflow is typically triggered.

The coaching field — converting scorecard into coaching tool

A scorecard that produces a number but no narrative has limited coaching value. Agents can see that they scored 72 on resolution, but without evaluator notes they cannot understand what specifically drove the score down or what a higher-scoring version of that interaction would have looked like.

Every section of the scorecard should include a comments and coaching field where the evaluator records:

  • What specific behavior or statement triggered the score on each criterion
  • The timestamp in the recording where the evaluator observed it
  • What a correct behavior would have looked like in this specific context
  • Any context that explains why the agent may have struggled with this criterion

Agents should have access to their completed scorecards, including the coaching fields. A score without notes is a judgment. A score with specific, timestamped observations tied to defined criteria is a coaching conversation. The difference in agent receptiveness and behavior change is significant.

Coaching fields also create a record that managers can review across evaluators to identify whether different evaluators are applying the same standard or whether one evaluator consistently leaves sparse notes while another writes detailed explanations. That discrepancy affects calibration regardless of whether the numeric scores differ.

Building for the channel

A voice scorecard evaluates what can be heard. A chat or email scorecard evaluates what can be read. These are different instruments, and treating them as equivalent leads to criteria that cannot be applied fairly across channels.

Voice-specific criteria include tone, pace, pronunciation, hold announcement and procedure, and background noise compliance. None of these translate directly to a chat interaction, where the customer cannot hear the agent at all. Chat-specific criteria include grammar and spelling accuracy, response time within SLA, appropriate use of links or resources, and managing concurrent chat volume without the conversation showing signs of context-switching errors.

Email QA adds criteria around completeness (was the customer's full question addressed), appropriate formality, and accuracy of attachments or links. Social media interactions introduce public-facing tone considerations that do not apply in private channels.

Building a QA program for an omnichannel contact center means maintaining separate scorecards calibrated to each channel's specific requirements, not adapting one voice scorecard to cover everything. Shared sections — like resolution accuracy and compliance disclosures — can use similar criteria across channels. Channel-specific sections need to be rewritten from scratch.

Testing and iterating the scorecard

A scorecard should be calibrated before deployment and reviewed regularly after. A scorecard released without calibration will generate inconsistent scores from the first day, and the inconsistency will erode agent trust in the entire QA program.

Pre-deployment calibration involves having two or more evaluators independently score the same set of recordings, then comparing scores criterion by criterion. Criteria where every evaluator agrees are either very clear or trivially easy — worth keeping if they matter, worth examining if they might not be measuring anything meaningful. Criteria generating constant disagreement are almost always poorly written and need to be revised before launch.

After deployment, run a pilot period — typically four to six weeks — during which scores are collected but not used for formal performance reviews. Use that period to identify:

  • Criteria where nearly every call scores the maximum — this may indicate the bar is too low or the criterion is not actually differentiating performance
  • Criteria where nearly every call scores zero — this may indicate the standard is unrealistic, or the criterion is capturing a training gap that needs to be addressed separately before it is scored
  • Weight distributions that, when analyzed against CSAT or first-call resolution data, do not correlate with actual customer outcomes — if your highest-weighted section has no predictive relationship to CSAT, either the weighting or the criteria need revision

Scorecard design is not a one-time event. As products change, regulations evolve, and your contact channel mix shifts, criteria and weights should be reviewed at least quarterly. A scorecard calibrated to last year's product lineup and last year's compliance environment is measuring the wrong things.

Frequently asked questions

How many criteria should a QA scorecard include? +
Most effective scorecards contain between 12 and 25 criteria across all sections. Fewer than 12 criteria tend to miss important behaviors or collapse distinct items into one criterion that becomes difficult to score consistently. More than 25 criteria create evaluator fatigue, extend evaluation time, and often include items that are redundant or marginally relevant. If your scorecard is growing beyond 25 criteria, review each item against the question: would a significant failure on this criterion change what the agent needs to do differently? Items that do not pass that test are candidates for removal.
Should QA scores be tied directly to agent compensation? +
Tying QA scores to compensation is common but requires a mature, calibrated program first. If evaluators are not scoring consistently and the scorecard has not been validated against actual outcomes, using those scores for pay decisions creates unfairness and erodes program credibility. Most programs run QA scores as a coaching tool for at least one to two calibration cycles before incorporating them formally into performance reviews or incentive plans. When they are used for compensation, they are typically weighted alongside other metrics — CSAT, handle time, attendance — rather than used as a standalone measure.
How many calls per agent should be evaluated each month? +
Common practice for voice queues ranges from 4 to 10 calls per agent per month, depending on team size, call volume, and the evaluation resource available. The sample must be large enough to be statistically meaningful — a single call per agent per month is too small a sample to trend accurately — but small enough that evaluators can complete reviews with sufficient depth and coaching notes. High-risk queues, new agents in their first 90 days, and agents on performance improvement plans are typically evaluated at higher sample rates than tenured agents in stable performance.
What is a calibration session and how often should they run? +
A calibration session is a meeting where evaluators — QA analysts, supervisors, and often a team lead or QA manager — independently score the same call recording, then compare scores and discuss criteria where they disagreed. The goal is not to reach consensus by committee but to identify where the scorecard language is ambiguous and agree on the correct interpretation so future scores are consistent. Most programs run calibration sessions weekly or bi-weekly, especially after any scorecard update or when new evaluators are onboarded. Calibration data — how much evaluator scores diverge — is itself a quality metric worth tracking over time.
Can AI replace human QA evaluators? +
AI-assisted QA tools can analyze 100% of calls rather than a sample, flag calls that match patterns associated with low scores or compliance risk, and automate scoring on criteria that are binary and clearly defined — such as whether a required phrase was spoken. However, nuanced criteria that require judgment about context, tone, or agent intent still benefit from human evaluation. The most common model is AI screening to identify calls that need human review, combined with human evaluation on the flagged calls and a representative sample of the rest. This extends QA coverage without proportionally increasing evaluator headcount.

A well-built scorecard does not operate in isolation. It feeds directly into QA calibration sessions that keep evaluator standards aligned, draws on call recording as the source material for every evaluation, and ultimately connects to CSAT measurement as the external validation that your internal quality standards are tracking what customers actually experience. Designing those three elements to work together — scorecard, recording, and outcome measurement — produces a QA program that improves performance rather than simply documenting it.

Related articles

CCaaS & Contact Center

Omnichannel Routing: How to Assign Voice, Chat, SMS, and Social Conversations

Omnichannel routing assigns customer interactions across every channel — voice, chat, SMS, email, and social — to the right agent based on skills, priority, load, and customer history. This guide explains how unified routing engines work, how concurrency differs across channels, and what to configure for consistent service levels.

CCaaS & Contact Center

SMS Character Limits, Encoding, and Message Segments Explained

A single emoji or unsupported character can change an SMS from one 160-character segment to two UCS-2 segments of 67 characters each — doubling your cost. This guide explains GSM-7, UCS-2, concatenation headers, segment counting, and how encoding decisions affect deliverability and billing.

Get Started

Quality Management for Contact Centers

Build scorecards, record calls, and run calibration sessions with integrated QA tools.