Most contact centers run a QA program that monitors and scores. The ones that actually get better add two things on top: practice sessions where agents rehearse before the next live call, and leaderboards that credit development instead of only raw output. This post maps that working loop onto the four surfaces under Quality in the Devotel Orbit dashboard — evaluations, a recording library, practice, and a leaderboard — and includes a sample rubric a supervisor can wire today.
Why agent QA programs plateau
The plateau usually shows up around the first or second quarter. The program launches, reviewers score a sample of conversations, coaching notes land, scores dip, a clinic happens, and then everything settles into a steady line for months.
The plateau has a structural cause, not a motivation problem. Three mechanics produce it:
- Surveillance without a remediation lane. The program measures quality, but fixing a weakness means having a bad conversation with a real customer first. A leaderboard built on that data punishes the agent for the flaw before anyone gave them a place to practice it.
- Volume-only leaderboards. When rankings sort on calls handled, the agents who need coaching most — the ones at the bottom — stop trusting the board. Public rankings on offenses ("lowest CSAT this week") breed concealment, and the dataset the QA team needs disappears.
- Fixed scoring weights forever. A weighted formula that never changes teaches agents which dimensions to game. They optimize the number, and the quality the number was a proxy for stays where it was.
A QA leaderboard works crediting development across the whole loop: detection (evaluations and recordings), remediation (practice), and a metric that blends those with live performance.
What to measure
Three inputs, each with its own dashboard surface in Devotel Orbit:
Evaluations. Reviewers build weighted scorecard forms and grade conversations; the evaluated agent can acknowledge or appeal, and a supervisor resolves the appeal. That appeal path matters more than reviewers expect — an agent can contest a score on the record, which keeps the data honest enough to rank people on. It lives at /quality/evaluations.
Recordings. Pulling the recordings behind the scores keeps sampling honest: any conversation in the library can be joined to its diarised transcript, its linked QA scorecard score, and a recording QC verdict. When an agent challenges a score, the reviewer opens /quality/recordings and looks at the source rather than arguing from memory.
Practice sessions. The remediation leg. Before the next live conversation, an agent rehearses the failing interaction type against an AI that plays the customer, and gets an automatic score plus coaching feedback. The roleplay is an in-app simulation — it never dials a real customer — so it is safe remediation and avoids the incentive to avoid difficult calls. This sits at /quality/practice.
The leaderboard then blends detection, remediation, and live performance at /quality/leaderboard. Skip any one leg and the program keeps measuring without improving.
Leaderboard mechanics used by top teams
Mechanics decide whether agents treat the board as a health metric or a surveillance report. Four patterns show up in programs that keep improving:
Rank within a queue, never across skill levels. The board is supervisor-scoped per queue. Compare like work with like work; a global top-10 list across queues mostly tells tenured, high-skill agents they are fine and new agents they are behind.
Credit development, not just output. A composite scoreboard with supervisor-tunable weights — the QA scorecard, calls handled, and CSAT — can move the needle on improvement. A newer agent climbing ranks beats a static veteran at the top; formula the supervisor adjusts in the scoring-rules editor so the team cannot game a fixed shape.
Include remediation effort. Grant partial credit for completed practice runs, or weight them beside live metrics. Agents will spend hours on hard-to-measure things when the board knows how to count them.
Rank AI and human agents on the same board. Where AI voice agents handle part of the load, the same scoring that judges them should feed the same table — the human-vs-AI gap closes when both are measured on one board instead of separate mirrors.
We wrote more about the scoring mechanics in Call-center agent-assist and quality scoring; the leaderboard is where the team sees those scores.
How Orbit's surfaces map each metric
Each component of a good program has a page in the Quality section of the Devotel Orbit dashboard. Open these in your workspace to build the loop; the links below go straight to the surface:
- [Evaluations](/quality/evaluations) — weighted scorecard forms; agents acknowledge or appeal, supervisors resolve. Detection.
- [Recording library](/quality/recordings) — every recorded conversation joined to its transcript, linked score, and QC verdict. Sampling honesty.
- [Practice Studio](/quality/practice) — agents rehearse against an AI customer with an auto-scored rubric; supervisors curate the scenario library. Remediation.
- [Performance leaderboard](/quality/leaderboard) — rankings computed on demand from QA scorecards, handled volume, and CSAT, with a supervisor-tuned scoring-weight editor. The engagement layer.
One tenant, one workspace, per-queue scope. Reviewer roles (owner, admin, supervisor) manage forms, weights, and the scenario library; agents see their own scores and sessions.
Sample leaderboard rubric
A concrete weighting a supervisor can start from, then tune:
| Signal | Weight | Source |
|---|---|---|
| Weekly QA scorecard average | 40% | Evaluations |
| Handled volume vs. team median | 20% | Queue metrics |
| CSAT (or post-contact survey) | 20% | Survey channel |
| Completed practice sessions | 10% | Practice Studio |
| Development credit (rank climb) | 10% | Leaderboard delta |
Run the development credit as a moving bonus: add points when an agent rises three or more ranks week over week, decaying after four weeks. New agents get a path to the top without discounting output quality. If you survey only two channels — QA and CSAT — move the practice weight into QA; the formula matters less than keeping it tunable and telling the team what changed when you re-tune.
Where the program goes next
Leaderboards work as the engagement layer of a QA loop; they are not the program on their own. For the remediation lane, see how the Practice Studio trains human agents with AI roleplay. For measuring the AI side of a blended team, see AI voice-agent quality assurance.
Frequently asked questions
Should the leaderboard be public to the whole team?
Per-queue scope makes the comparison fair, and the appeal path keeps the data honest. Most teams show rankings to the team but keep the full scorecard breakdowns reviewer-only.
How often should scoring weights change?
Rarely, and publicly. A quarterly re-tune with an announced change log keeps agents from gaming a static formula without making the goalposts move weekly.
Do practice sessions really belong in the ranking?
Yes, at a small weight. A 10% credit for completed runs moves remediation spend from invisible to counted, which changes whether agents actually do it.
Is the leaderboard an add-on module?
No. The ranking is computed on demand from data the Quality suite already captures — evaluations, handled volume, CSAT — so the surface is part of the QA workspace, not a separate gamification purchase.