A QA program without a scorecard is just opinions. Different evaluators notice different things, weight them differently, and deliver feedback that agents cannot act on consistently. The scorecard is the instrument that converts a subjective listening experience into a structured, repeatable measurement — one that means the same thing whether a supervisor in one city evaluates a call or an evaluator in another city evaluates a different call the same morning.
Building that instrument well requires more than listing behaviors you want to see. It requires deciding which behaviors matter most, how to describe them so that any evaluator would score them the same way, which violations are severe enough to fail an entire call, and how to make the scorecard function as a coaching tool rather than just a report card.
A call center QA scorecard is a structured evaluation instrument that breaks agent performance into scored sections and criteria, enabling consistent, quantitative assessment of call quality across a team. Each section groups related behaviors, each criterion is scored against a defined standard, and the weighted total produces a call score that can be trended, benchmarked, and used to direct coaching.
Scorecard sections and their purpose
Most QA scorecards organize criteria into five or six sections that map to the natural arc of a customer interaction. Grouping related criteria serves two purposes: it makes scoring easier to follow, and it gives evaluators and agents a clear picture of which phase of the call had problems. A score breakdown by section is far more actionable than a single total.
Opening and greeting
The opening section covers the first twenty to thirty seconds of the call: did the agent answer within the required number of rings, introduce themselves and the company correctly, confirm they are speaking with the right person, and set a tone that is warm and appropriately paced? A poor opening can undermine the entire interaction before the customer has stated a single word about their issue.
Criteria in this section typically evaluate identity verification at the point of greeting (where required), use of the customer's name, and absence of behaviors that signal disengagement — rushed delivery, mispronounced company name, or failure to acknowledge the caller.
Needs identification and discovery
This section evaluates whether the agent asked enough of the right questions to fully understand what the customer needed before moving to resolution. Active listening markers — paraphrasing, acknowledgment statements, open-ended probing questions — belong here. So does the agent's ability to distinguish between the customer's stated request and their underlying need, which are not always the same thing.
Agents who skip this phase and jump to resolution based on incomplete information generate repeat contacts and lower first-call resolution rates. The scorecard makes that visible by scoring discovery as its own section rather than folding it into resolution.
Resolution and product knowledge
Resolution criteria assess whether the agent provided an accurate, complete answer to the customer's issue. This is typically the highest-weighted non-compliance section because it maps most directly to the reason the customer called. Criteria include: was the information provided factually correct, was the resolution complete or partial, did the agent take appropriate ownership rather than deflecting, and was the first-call resolution achieved or was a follow-up necessary?
Product knowledge is evaluated here as well. An agent who provides a confident but incorrect answer scores worse than one who acknowledges uncertainty and escalates appropriately, because the former creates downstream problems the customer will have to call back about.
Compliance and regulatory
Compliance criteria cover legally or contractually required behaviors: reading required disclosures, completing mandated verification steps, following script requirements where they apply, and avoiding prohibited statements. In regulated industries — financial services, healthcare, collections, insurance — this section carries the most weight and is most likely to contain auto-fail items.
Even in less regulated environments, compliance criteria often include internal policies: not making unauthorized commitments, not discussing competitor pricing in certain ways, not offering discounts beyond agent authority. These protect the organization even when no external regulator is watching.
Call handling mechanics
This section covers the procedural elements of managing a call: placing the customer on hold correctly (announcing the hold, giving a time estimate, checking back if the hold runs long), executing a transfer with the appropriate warm handoff or cold-transfer procedure, completing after-call work within the target window, and logging the call accurately. These behaviors are largely binary — either the agent followed the procedure or did not — which makes them among the easiest to score consistently.
Closing
The closing section evaluates whether the agent brought the call to a proper conclusion: summarizing what was resolved or agreed, confirming the customer's next steps, asking whether anything else is needed, and ending the call with appropriate acknowledgment. A weak closing leaves customers uncertain about what happens next and contributes to repeat contacts. Evaluators look for whether the agent owned the close or let the call trail off.
Weighting — not all criteria are equal
Once you have your sections and criteria, you need to assign weights. Weighting reflects your organization's priorities: a criteria set where compliance carries the same weight as hold procedure sends the wrong signal about what matters.
A common starting framework — illustrative only, to be adjusted to your business context — distributes section weights approximately as follows:
- Compliance and regulatory: 30% — highest weight because failures here carry legal, financial, or reputational consequences
- Resolution and product knowledge: 35% — weighted above communication because solving the customer's issue is the core purpose of the interaction
- Communication and soft skills: 20% — important for customer experience but secondary to whether the issue was actually resolved correctly
- Call handling mechanics: 15% — procedural adherence matters but rarely drives CSAT by itself
Within sections, individual criteria can carry different weights. A single item covering mandatory regulatory disclosure might represent 20 of the 30 compliance points, while a greeting-format criterion represents 2 of the opening's 10 points. Weight by business outcome: what does a failure on this criterion actually cost the customer, the organization, or the regulatory relationship?
Resist the temptation to weight every criterion equally because it is simpler. Equal weighting implies that missing the company's greeting script is as important as failing to verify caller identity before discussing account details. It is not, and a flat weighting system will not surface that distinction in your scores.
Observable versus subjective criteria
The single most common reason QA programs fail to produce consistent scores is criteria written around impressions rather than observable behaviors. The difference determines whether two evaluators listening to the same call would score it the same way.
A subjective criterion sounds like: "Agent was professional and courteous." There is no way to score this without applying personal interpretation. One evaluator's "professional" is another's "cold." The result is evaluator disagreement, which undermines calibration and makes the scorecard impossible to defend when an agent contests a score.
An observable criterion sounds like: "Agent verified caller name and account number before discussing account details." Either the agent did this or did not. The evaluator does not need to interpret anything. Observable criteria describe specific, audible behaviors that any evaluator can identify by listening to the recording.
Converting subjective criteria to observable ones requires asking: what would I actually hear if the agent did this correctly? Write that behavior down in specific terms. "Agent used empathy statement after customer expressed frustration" is more observable than "Agent demonstrated empathy." "Agent read the required fraud disclosure verbatim before processing the transaction" is more observable than "Agent followed compliance procedures."
Observable criteria reduce evaluator disagreement, make calibration sessions shorter and more productive, and give agents unambiguous feedback they can act on.
Auto-fail criteria
Auto-fail criteria are behaviors severe enough that a single instance fails the entire call, regardless of how well the agent performed on every other criterion. Auto-fails exist because some behaviors cannot be partially offset by good performance elsewhere. A call on which the agent misrepresented a product feature is not an 80-point call because the hold procedure was excellent.
Common auto-fail categories across industries:
- Failure to verify caller identity before discussing or modifying account information — a direct security and compliance failure
- Making a prohibited promise — committing the organization to something outside agent authority or policy (a specific refund amount, a callback on a date the system cannot support)
- Missing a legally required disclosure — in regulated industries, omitting a mandated disclosure can carry regulatory consequences independent of whether the customer noticed
- Misrepresenting a product or service — providing materially false information about features, pricing, or terms
- Unprofessional conduct — abusive, discriminatory, or retaliatory behavior toward the customer
Auto-fail criteria must be documented explicitly, separately from weighted scoring criteria, and reviewed with all agents as part of onboarding and calibration. Every agent should be able to name your auto-fail behaviors without looking them up. When an auto-fail occurs, the call receives a zero or a designated failing score, and a separate coaching and documentation workflow is typically triggered.
The coaching field — converting scorecard into coaching tool
A scorecard that produces a number but no narrative has limited coaching value. Agents can see that they scored 72 on resolution, but without evaluator notes they cannot understand what specifically drove the score down or what a higher-scoring version of that interaction would have looked like.
Every section of the scorecard should include a comments and coaching field where the evaluator records:
- What specific behavior or statement triggered the score on each criterion
- The timestamp in the recording where the evaluator observed it
- What a correct behavior would have looked like in this specific context
- Any context that explains why the agent may have struggled with this criterion
Agents should have access to their completed scorecards, including the coaching fields. A score without notes is a judgment. A score with specific, timestamped observations tied to defined criteria is a coaching conversation. The difference in agent receptiveness and behavior change is significant.
Coaching fields also create a record that managers can review across evaluators to identify whether different evaluators are applying the same standard or whether one evaluator consistently leaves sparse notes while another writes detailed explanations. That discrepancy affects calibration regardless of whether the numeric scores differ.
Building for the channel
A voice scorecard evaluates what can be heard. A chat or email scorecard evaluates what can be read. These are different instruments, and treating them as equivalent leads to criteria that cannot be applied fairly across channels.
Voice-specific criteria include tone, pace, pronunciation, hold announcement and procedure, and background noise compliance. None of these translate directly to a chat interaction, where the customer cannot hear the agent at all. Chat-specific criteria include grammar and spelling accuracy, response time within SLA, appropriate use of links or resources, and managing concurrent chat volume without the conversation showing signs of context-switching errors.
Email QA adds criteria around completeness (was the customer's full question addressed), appropriate formality, and accuracy of attachments or links. Social media interactions introduce public-facing tone considerations that do not apply in private channels.
Building a QA program for an omnichannel contact center means maintaining separate scorecards calibrated to each channel's specific requirements, not adapting one voice scorecard to cover everything. Shared sections — like resolution accuracy and compliance disclosures — can use similar criteria across channels. Channel-specific sections need to be rewritten from scratch.
Testing and iterating the scorecard
A scorecard should be calibrated before deployment and reviewed regularly after. A scorecard released without calibration will generate inconsistent scores from the first day, and the inconsistency will erode agent trust in the entire QA program.
Pre-deployment calibration involves having two or more evaluators independently score the same set of recordings, then comparing scores criterion by criterion. Criteria where every evaluator agrees are either very clear or trivially easy — worth keeping if they matter, worth examining if they might not be measuring anything meaningful. Criteria generating constant disagreement are almost always poorly written and need to be revised before launch.
After deployment, run a pilot period — typically four to six weeks — during which scores are collected but not used for formal performance reviews. Use that period to identify:
- Criteria where nearly every call scores the maximum — this may indicate the bar is too low or the criterion is not actually differentiating performance
- Criteria where nearly every call scores zero — this may indicate the standard is unrealistic, or the criterion is capturing a training gap that needs to be addressed separately before it is scored
- Weight distributions that, when analyzed against CSAT or first-call resolution data, do not correlate with actual customer outcomes — if your highest-weighted section has no predictive relationship to CSAT, either the weighting or the criteria need revision
Scorecard design is not a one-time event. As products change, regulations evolve, and your contact channel mix shifts, criteria and weights should be reviewed at least quarterly. A scorecard calibrated to last year's product lineup and last year's compliance environment is measuring the wrong things.
Frequently asked questions
A well-built scorecard does not operate in isolation. It feeds directly into QA calibration sessions that keep evaluator standards aligned, draws on call recording as the source material for every evaluation, and ultimately connects to CSAT measurement as the external validation that your internal quality standards are tracking what customers actually experience. Designing those three elements to work together — scorecard, recording, and outcome measurement — produces a QA program that improves performance rather than simply documenting it.