Skip to main content
Back to blog

The supervisor coaching queue, explained — threshold triage to acknowledged note

Devotel Orbit's Inbox → Coaching page turns LLM-judged rubric scores into a supervisor worklist — the low-scoring conversations worth a human review, lowest first, with a coaching note that closes only when the agent acknowledges it. Here is the queue math, the failure-mode frame, and the cadence that keeps it honest.

Orbit Editorial Team

Most quality programs have a score and a wish. The score arrives from the rubric judge, the wish is that a supervisor somehow finds the bad conversations, reads them, and says something that changes the next hundred. Devotel Orbit's Inbox → Coaching page removes the wish — it is a supervisor worklist built directly from the scored rubrics: conversations that fell below your threshold, sorted worst-first, each one one click from a coaching note the agent must acknowledge.

The competitive frame is the same one the gamification posts make against NICE, Genesys, and Talkdesk: premium QM suites sell the queue as part of a gamification bundle. Here it is a page in the inbox, driven by the same rubric scores your QA program already produces, with no leaderboard required to make it work.

What a supervisor coaching queue is

Open Inbox → Coaching and you get a single triage table. Each row is one conversation the QM judge scored against your rubrics and found below the line: the contact, a channel marker (inbox or voice), the average score as a percentage, the failed rubric count as a fraction (2/5 — two of five rubrics failed), and when it was last scored. Rows arrive lowest-score-first, so the worst of the window is always at the top.

The row is a doorway, not a destination. Click it and you land inside the conversation itself — the thread on the inbox pane for digital channels, the call surface for voice — where you read what actually happened before writing anything. The queue's job is to make sure the right conversations get read; the review still belongs to the human.

Empty is a valid state. A workspace with no scored conversations yet, or one where everything passed, shows a clean "no low-scoring conversations" state — not an error. A queue that never empties because nobody writes notes is a program problem; a queue that is empty because nothing failed is the goal.

The threshold arithmetic

"Low scoring" is not a platform verdict — it is a comparison against a number the supervisor sets on the page. A conversation enters the queue when two things are both true:

  1. The average rubric score falls below the threshold. Every rubric the judge ran on the conversation produces a confidence value between 0 and 1; the average of those values must be strictly below the configured threshold to flag the conversation.
  2. At least one rubric failed. A conversation cannot ride into the queue on a low average alone — there must be a named failure to coach against. This keeps the table free of statistically odd but genuinely-passed conversations.

The page offers four threshold positions: Strict (average under 30%), Default (under 50%), Lenient (under 70%), and All scored (the threshold check off — every conversation with at least one failed rubric, still sorted worst-first). Most teams run the default and tighten when sampling volume is high, loosen when the queue runs dry. The reappraisal is cheap because the filter is a page control, not a config deployment.

Two arithmetic properties follow from the construction. First, sortedness: worst-first ordering is stable down the list because ties break on the conversation id, so a supervisor paging through "Load more" sees a deterministic sequence. Second, monotonicity: raising the threshold strictly widens the queue, so a threshold change never hides a conversation that was previously visible.

The failure-mode frame

A score says that a conversation went badly; the rubric breakdown says how. Each failed rubric the judge flagged becomes a named failure mode — "skipped identity verification," "promised a feature we do not ship," "closed without resolving the customer's question" — and those names accumulate across the queue into the top failure modes list on the agent's scorecard.

That frame is what a coaching note attaches against. When a supervisor writes a note, they reference the specific failed rubrics (up to ten per note), so the feedback is evidence pinned to the review, not a general admonishment. The note body is free-form text with an optional suggested action ("re-run the verification walkthrough before your next shift"), and it stays on the conversation — visible in the coaching history of the thread, newest first, with the agent's acknowledgement stamp once they read it.

The taxonomy discipline matters at the queue level. If every note cites the rubric it targets, the failure-mode counts stay honest, and the per-agent scorecard — per-rubric pass rates sorted from weakest up — stays the reliable coaching agenda for the weekly one-on-one.

Supervisor drill-in

The queue page is role-gated: owners, admins, and the supervisor seat see it; agents never see the queue — their view of the loop is their own coaching notes and their own scorecard. Two controls shape what a supervisor drills into:

  • Channel filter. All, Inbox, or Voice. The filter exists because the QM judge scores both digital threads and calls, and the same queue surfaces both — a voice row deep-links into the call surface rather than the inbox pane.
  • Threshold selector. The four positions above, applied to the whole queue.

Pagination is cursor-based: twenty rows per page, with a "Load more" control that walks the list in stable worst-first order. The resets are explicit — changing the channel or the threshold clears the accumulated pages and re-fetches from the top, so a filter swap never strands stale rows in the list.

The queue also refuses to pretend: a fetch error surfaces as a retryable block with the message the API returned, and the empty-after-filter state offers a one-click reset back to the defaults. A supervisor triage tool that hides its own failures would erode the trust the whole loop depends on.

The cadence: weekly processing against a threshold-driven queue

The queue is threshold-driven — entries appear as the judge scores newly closed conversations below the line — but the processing of it is a supervisor habit, and the sustainable one is weekly:

  1. Review the queue worst-first. Read the thread before writing the note. Queue rows tell you which rubric failed, not why — the why lives in the transcript.
  2. Write one note per flagged conversation, citing the failed rubrics. Short, specific, weekly notes outperform the quarterly write-up, and the acknowledgement stamp gives you a hard read receipt per agent.
  3. Check the scorecard, not just the queue. The per-agent scorecard carries the top failure modes and a daily score trend over up to 90 days; a queue that keeps surfacing the same failure mode from the same agent is a coaching-plan candidate, not another note.
  4. Measure whether it worked. Once a note lands, the before/after read on the agent's digital KPIs — first-response time, resolution rate, and CSAT — tells you whether the coaching actually moved the numbers, or whether you are coaching the wrong thing.

If your QA program needs the queue to reach people who live outside the dashboard, the tenant's existing notification surfaces carry it — the same event and webhook plumbing that routes your other workspace events can push digest-shaped signals to email or Slack on your own schedule, while the queue itself remains the system of record. The loop closes in the app; the reminders can live wherever your supervisors already are.

Tenant-owned QA posture

Nothing in this loop is a platform mandate. Every control is owned by the tenant, from inside the tenant's own organization settings:

  • The threshold is yours. The page selector above is a workspace control; the API accepts the same parameter, so scripted triage respects whatever your supervisors actually run.
  • The notes are yours. Coaching notes are authored and readable by owner/admin/supervisor, addressed to the specific agent on the conversation, and the addressed agent is the only seat that can acknowledge them. Neither the queue nor the notes ever leave your workspace — nothing customer-facing is generated.
  • The sampling feeding the queue is yours. If the queue runs dry at a threshold you expect traffic against, the knob to turn is the QA sampler's assignment quota, not the platform's defaults.

That posture generalizes beyond the queue — it is the same tenant-owned stance the compliance guides take on quiet hours and consent: the platform provides the workspace, the program belongs to you.

Frequently asked questions

Does the queue treat inbox and voice conversations the same?

Yes — the same threshold and the same worst-first ordering apply to both, and the channel filter is a view control, not a scoring difference. Voice rows deep-link into the call surface; inbox rows deep-link into the thread pane. The remaining asymmetry is on the coaching side: digital feedback stays an acknowledged note on the thread, while multi-week voice remediation runs through the coaching-plan workspace.

Where do the scores come from?

The queue reads the LLM-judged rubric outcomes your QA program produces when scored conversations close. A rubric registry without assigned sampling produces an empty queue, so the scoring feed — the sampler and the rubric set, configured in your quality settings — is the upstream dependency to keep healthy.

What is the lifecycle of a coaching note?

A supervisor posts the note on a conversation, optionally citing failed rubrics and a suggested action. The note is visible to supervisors and to the addressed agent — and only them — in the conversation's coaching history, newest first. It stays open until the addressed agent acknowledges it with one action; the acknowledgement stamp is then visible on the note back to every supervisor reading the history. Unacknowledged notes are the loop's open work.

Reference: the queue and the leaderboard

The coaching queue and the gamification leaderboard answer different questions. The queue's question is "which conversations need a human now"; the leaderboard's is "how do we keep agents engaged with quality over quarters." They are siblings, not substitutes — and the leaderboard argument (development-crediting mechanics instead of raw-output ranking) is exactly the one in Contact-center leaderboards that work, while the intervention side of the same loop — plans with goals, owners, and measured outcomes — is covered in Coaching plans that close the loop.

For the product surface this post explains, see the dashboard at Inbox → Coaching (and its agent-facing mirror at /me/scorecard), or the endpoint walkthrough in the inbox coaching docs. The queue is where the program becomes daily; the leaderboard is where it becomes a culture; the plans are where it becomes measurable. You want all three anchored on the same scores — that is the whole posture.

The supervisor coaching queue, explained — threshold triage to acknowledged note — Orbit by Devotel